← Rohit Girdhar
Publications
2026
Muse Image and Video
In
Meta AI Blog
, 2026
2026
Fast, High-Fidelity Video Editing
In
Meta AI & Instagram Edits
, 2026
2026
Human detectors are surprisingly powerful reward models
In
arXiv
, 2026
2025
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
In
CVPR
, 2026
2025
Diffusion Autoencoders are Scalable Image Tokenizers
In
arXiv
, 2025
2025
LLMs can see and hear without any training
In
ICML
, 2025
2024
MotiF: Making Text Count in Image Animation with Motion Focal Loss
In
CVPR
, 2025
2024
Movie Gen: A Cast of Media Foundation Models
In
arXiv
, 2024
2024
The Llama 3 Herd of Models
In
arXiv
, 2024
2024
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
In
CVPR
, 2024
2024
InstanceDiffusion: Instance-level Control for Image Generation
In
CVPR
, 2024
2023
Generating Illustrated Instructions
In
CVPR
, 2024
2023
Motion-Conditioned Image Animation for Video Editing
In
arXiv
, 2023
2023
VideoCutLER: Surprisingly Simple Unsupervised Video Instance Segmentation
In
CVPR
, 2024
2023
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
In
ECCV
, 2024
2023
ImageBind: One Embedding Space To Bind Them All
In
CVPR
, 2023
(Highlighted Presentation)
2023
The effectiveness of MAE pre-pretraining for billion-scale pretraining
In
ICCV
, 2023
2023
CutLER: Cut and Learn for Unsupervised Object Detection and Instance Segmentation
In
CVPR
, 2023
2023
HierVL: Learning Hierarchical Video-Language Embeddings
In
CVPR
, 2023
(Highlighted Presentation)
2022
Learning Video Representations from Large Language Models
In
CVPR
, 2023
(Highlighted Presentation)
2022
OmniMAE: Single Model Masked Pretraining on Images and Videos
In
CVPR
, 2023
2022
Omnivore: A Single Model for Many Visual Modalities
In
CVPR
, 2022
(Oral Presentation)
2022
Ego4D: Around the World in 3,000 Hours of Egocentric Video
In
CVPR
, 2022
(Best paper finalist)
2022
Detecting Twenty-thousand Classes using Image-level Supervision
In
ECCV
, 2022
2021
Mask2Former for Video Instance Segmentation
In
arXiv
, 2021
2021
Masked-attention Mask Transformer for Universal Image Segmentation
In
CVPR
, 2022
2021
3DETR: An End-to-End Transformer Model for 3D Object Detection
In
ICCV
, 2021
(Oral Presentation)
2021
Anticipative Video Transformer
In
ICCV
, 2021
2021
3D Spatial Recognition without Spatially Labeled 3D
In
CVPR
, 2021
2021
Self-Supervised Pretraining of 3D Features on any Point-Cloud
In
CVPR
, 2021
2021
Physical Reasoning Using Dynamics Aware Embeddings
In
ICML Workshops
, 2021
2020
Forward Prediction for Physical Reasoning
In
ICML Workshops
, 2021
2019
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
In
ICLR
, 2020
(Oral Presentation)
2019
MetaPix: Few-Shot Video Retargeting
In
ICLR
, 2020
2019
DistInit: Learning Video Representations Without a Single Labeled Video
In
ICCV
, 2019
2018
Video Action Transformer Network
In
CVPR
, 2019
(Oral Presentation)
2017
Detect-and-Track: Efficient Pose Estimation in Videos
In
CVPR
, 2018
2017
Attentional Pooling for Action Recognition
In
NeurIPS
, 2017
2017
ActionVLAD: Learning spatio-temporal aggregation for action classification
In
CVPR
, 2017
2016
Binge Watching: Scaling Affordance Learning from Sitcoms
In
CVPR
, 2017
(Spotlight Presentation)
2016
Learning a Predictable and Generative Vector Representation for Objects
In
ECCV
, 2016
(Spotlight Presentation)