rohitgirdhar.github.io
Photo of Rohit Girdhar

Rohit Girdhar

Research Scientist · AMI Labs

I am a Research Scientist at AMI Labs. My current research focuses on multimodal understanding, generation and world modeling. I obtained a PhD from Carnegie Mellon University (here’s a link to my dissertation), where I worked on learning from and understanding videos. I was previously part of the Meta Superintelligence Labs (MSL) and Facebook AI Research (FAIR) at Meta, and have spent time at DeepMind, Adobe and Facebook as an intern. See here for a formal bio.

  1. 2026 — Present

    AMI Labs

    Research Scientist · New York

    Building world models that understand real-world sensor data. Research on multimodal understanding, generation and world modeling.

  2. 2019 — 2026

    Meta FAIR, GenAI & MSL

    Research Scientist · New York

    Core contributor to Meta’s media generation and multimodal efforts, including Movie Gen, Muse Image, Emu Video, ImageBind and Llama 3. Earlier at FAIR, worked on self-supervised, multimodal and omnivorous representations for video understanding.

    1. 2026
      Muse Image and Video In Meta AI Blog, 2026
    2. 2026
      Fast, High-Fidelity Video Editing In Meta AI & Instagram Edits, 2026
    3. 2026
    4. 2025
    5. 2025
    6. 2025
    7. 2024
    8. 2024
    9. 2024
    10. 2024
    11. 2024
    12. 2023
    13. 2023
    14. 2023
    15. 2023
    16. 2023
      ImageBind: One Embedding Space To Bind Them All In CVPR, 2023 (Highlighted Presentation)
    17. 2023
    18. 2023
    19. 2023
      HierVL: Learning Hierarchical Video-Language Embeddings In CVPR, 2023 (Highlighted Presentation)
    20. 2022
      Learning Video Representations from Large Language Models In CVPR, 2023 (Highlighted Presentation)
    21. 2022
    22. 2022
      Omnivore: A Single Model for Many Visual Modalities In CVPR, 2022 (Oral Presentation)
    23. 2022
    24. 2022
    25. 2021
    26. 2021
    27. 2021
    28. 2021
    29. 2021
    30. 2021
    31. 2021
    32. 2020
  3. Summer 2018

    DeepMind

    Research Scientist Intern · London

    Worked with Andrew Zisserman, João Carreira and Carl Doersch on attention-based architectures for recognizing and localizing human actions in video.

  4. Summer 2017

    Facebook AI Research

    Research Scientist Intern · Menlo Park

    Worked with Lorenzo Torresani, Georgia Gkioxari and Du Tran on detecting and tracking people for keypoint estimation in video.

  5. Summer 2016

    Adobe Research

    Research Scientist Intern · San Francisco

    Worked with Josef Sivic and Bryan Russell on learnable spatio-temporal aggregation for action classification.

  6. 2014 — 2019

    Carnegie Mellon University The Robotics Institute

    PhD in Robotics · Pittsburgh

    PhD with Deva Ramanan on learning from and understanding videos; MS with Martial Hebert, Abhinav Gupta, Kris Kitani and David Fouhey as a Siebel Scholar.

  7. Summer 2013

    Facebook

    Software Engineering Intern · Menlo Park

    First stint at Facebook, working on Graph Search.

  8. 2010 — 2014

    IIIT Hyderabad

    B.Tech in Computer Science · Hyderabad, India

    Undergraduate research with C. V. Jawahar at CVIT on large-scale image retrieval and computer vision.