Meta Superintelligence Labs enters the picture with Muse
Meta Superintelligence Labs launches Muse Image, an agentic AI model ranking second on the Arena, alongside a preview of its Muse Video generator.
Internally codenamed Mango, Muse Image is the first image model from Meta Superintelligence Labs, the AI division Mark Zuckerberg restructured at great expense to catch up, led by Alexandr Wang, following the Muse Spark language model released in April. Its uniqueness lies less in its rendering than in its operation: it acts as an agent. Before producing an image, it reasons, queries the web for context, writes and executes code, then rereads and corrects its own drafts—a self-revision Meta says it saw emerge on its own through reinforcement learning rather than having programmed it. It starts from user photos or tagged accounts to restore an old family snapshot, change a hairstyle, transform a face into a clay figurine, or create a photo booth-style portrait. This reliance on code also allows it to produce a truly scannable QR code, accurate graphics, or legible text within the image.
On the Arena, the reception leans positive: Muse Image ranks second in both text-to-image and image editing, behind only OpenAI's GPT Image 2. Criticism is directed elsewhere, at the Instagram reference. Via an @-mention, the model can adopt the features of a third-party public account: the function is active by default, only deactivatable via opt-out, and the person concerned is not notified. The Verge, which raised the issue, WIRED, and Digital Trends note that already generated images are not deleted after deactivation and that the setting was not yet visible everywhere at launch. The invisible Content Seal watermark, present on every output, attests to the AI origin but provides no control over what has already been produced.
Muse Video, on the other hand, is only previewed. Built on the same pre-training foundation, it generates image and sound within a single model, without a separate track, and ranks third in human preference for text-to-video, behind Google's Gemini Omni Flash and ByteDance's Seedance. Meta says it is competitive in prompt following and temporal consistency, and acknowledges two blind spots: audio-video synchronization and rapid movements.
Deployment begins in the United States within Meta AI, with about thirty effects for Instagram stories and image generation in WhatsApp, Facebook, and Messenger to follow. Standard use is free up to a threshold, then switches to subscription; advertiser access via Advantage+ is planned within a few weeks. These in-house models are intended to reduce the company's dependence on third parties like Midjourney or Black Forest Labs, with the opening to external developers not yet decided. This bet on image and video contrasts with a sector now focused on code generation.