Muse Image and Video
In Meta AI Blog, 2026

Abstract
Muse Image and Muse Video are the first media generation models developed by Meta Superintelligence Labs (MSL). Rather than directly mapping text to pixels, Muse Image operates as an autonomous agent: it invokes search and coding tools to ground its generations in factual, real-time information and to render accurate plots, figures and text, self-refines its own outputs, and improves further by scaling test-time compute — exhibiting an approximately log-linear relationship between inference-time reasoning and image quality. Beyond text-to-image generation, the model performs precise instruction-following edits over multiple turns, and composes people, objects, clothing, styles and environments from many input reference images. Muse Video generates videos with native audio, with competitive prompt adherence, visual fidelity and temporal consistency. At launch, Muse Image ranked #2 on the Arena human-preference leaderboards for text-to-image, single-image editing, and multi-image editing, and Muse Video ranked #3 for text-to-video. The models power media generation in the Meta AI app, meta.ai, Instagram and WhatsApp. I drove key efforts in multi-image editing.