Publications

All publications

A selection of research across multimodal learning, video understanding, efficient model design, and machine learning systems.

Research

Publications

For the complete and most up-to-date record, visit Google Scholar.

Google Scholar ↗
2026
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
arXiv preprint
Emad Bahrami, Olga Zatsarynna, Parth Pathak, Sunando Sengupta, Jürgen Gall, Mohsen Fayyaz
2025
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
NeurIPS
Ali Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk, Mohsen Fayyaz
2025
DeltaLLM: Compress LLMs with Low-Rank Deltas between Shared Weights
arXiv preprint
Liana Mikaelyan, Ayyoob Imani, Mathew Salvaris, Parth Pathak, Mohsen Fayyaz
2024
Occlusion Handling in 3D Human Pose Estimation with Perturbed Positional Encoding
ECCV
Niloofar Azizi, Mohsen Fayyaz, Horst Bischof
2022
TaylorSwiftNet: Taylor Driven Temporal Modeling for Swift Future Frame Prediction
BMVC
Saber Pourheydari*, Emad Bahrami*, Mohsen Fayyaz*, Gianpiero Francesca, Mehdi Noroozi, Jürgen Gall
* Equal contribution
2022
Adaptive Token Sampling for Efficient Vision Transformers
ECCV — Oral Presentation (Top 3%)
Mohsen Fayyaz*, Soroush Abbasi Koohpayegani*, Farnoush Rezaei Jafari*, Sunando Sengupta, Hamid Reza Vaezi Joze, Eric Sommerlade, Hamed Pirsiavash, Jürgen Gall
* Equal contribution
2022
Fast Weakly Supervised Action Segmentation Using Mutual Consistency
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
Yaser Souri*, Mohsen Fayyaz*, Luca Minciullo, Gianpiero Francesca, Jürgen Gall
* Equal contribution
2021
3D CNNs with Adaptive Temporal Feature Resolutions
CVPR
Mohsen Fayyaz*, Emad Bahrami*, Ali Diba, Mehdi Noroozi, Ehsan Adeli, Luc Van Gool, Jürgen Gall
* Equal contribution
2021
Long Short View Feature Decomposition via Contrastive Video Representation Learning
ICCV
Nadine Behrmann, Mohsen Fayyaz, Jürgen Gall, Mehdi Noroozi
2020
Large Scale Holistic Video Understanding
ECCV — Spotlight (Top 3%)
Ali Diba*†, Mohsen Fayyaz*†, Vivek Sharma*†, Manohar Paluri, Jürgen Gall, Rainer Stiefelhagen, Luc Van Gool
* Equal contribution · † Listed in alphabetical order
2020
SCT: Set Constrained Temporal Transformer for Set Supervised Action Segmentation
CVPR
Mohsen Fayyaz, Jürgen Gall
2018
AVID: Adversarial Visual Irregularity Detection
ACCV
Mohammad Sabokrou*, Masoud Pourreza*, Mohsen Fayyaz*, Rahim Entezari, Mahmood Fathy, Jürgen Gall, Ehsan Adeli
* Equal contribution
2018
Spatio-Temporal Channel Correlation Networks for Action Classification
ECCV
Ali Diba*, Mohsen Fayyaz*, Vivek Sharma, M. Mahdi Arzani, Rahman Yousefzadeh, Jürgen Gall, Luc Van Gool
* Equal contribution
2018
Temporal 3D ConvNets by Temporal Transition Layer
CVPR Workshop on Brave New Ideas in Video Understanding
Ali Diba*, Mohsen Fayyaz*, Vivek Sharma, A. Karami, M. Arzani, Rahman Yousefzadeh, Luc Van Gool
* Equal contribution
2018
Deep-anomaly: Fully Convolutional Neural Network for Fast Anomaly Detection in Crowded Scenes
Computer Vision and Image Understanding
Mohammad Sabokrou*, Mohsen Fayyaz*, Mahmood Fathy, Zahra Moayed, Reinhard Klette
* Equal contribution
2018
Towards Principled Design of Deep Convolutional Networks: Introducing SimpNet
arXiv preprint
Seyyed Hossein Hasanpour, Mohammad Rouhani, Mohsen Fayyaz, Mohammad Sabokrou, Ehsan Adeli
2017
Deep-cascade: Cascading 3D Deep Neural Networks for Fast Anomaly Detection and Localization in Crowded Scenes
IEEE Transactions on Image Processing
Mohammad Sabokrou, Mohsen Fayyaz, Mahmood Fathy, Reinhard Klette
2016
STFCN: Spatio-Temporal Fully Convolutional Neural Network for Semantic Segmentation of Street Scenes
ACCV Workshop
Mohsen Fayyaz, Mohammad Sabokrou, Mohammad Hajizadeh, Mahmood Fathy, Fay Huang, Reinhard Klette