Meta releases Movie Gen video and audio foundation models
A 30B-parameter video model paired with a 13B-parameter audio model produced clips up to 16 seconds with synchronised sound; Meta called it research, not a product.
- Models & capabilities
- Notable
Meta announced Movie Gen, a set of foundation models for generating and editing video and audio from text. The video component, a 30-billion-parameter transformer, produced high-definition clips up to 16 seconds long from a text description, handling object motion, camera movement and interactions between subjects. A separate 13-billion-parameter audio model generated ambient sound, effects and music synchronised to the video, up to 45 seconds long.
Beyond straight text-to-video generation, Meta demonstrated two further capabilities: personalisation, in which a photo of a person was combined with a text prompt to produce a video that preserved their likeness and motion, and precise editing, in which an existing video could have elements added or removed, or its background and style changed, while the rest of the footage was left intact. Meta said human evaluators preferred Movie Gen’s outputs to those of competing systems across all four capabilities, without publishing the underlying win-rate figures in the announcement.
Movie Gen combined internet-scale pretraining with licensed and proprietary video data; Meta did not disclose the training dataset in detail. The release was explicitly framed as research rather than a product: no public access, API or model weights accompanied the announcement, and Meta said it was working towards “a potential future release” in collaboration with film-industry creators.
The announcement placed Meta alongside OpenAI’s unreleased Sora and Adobe’s Firefly video tools in a wave of 2024 demonstrations that showed video generation had moved from short, artefact-heavy clips towards coherent minute-scale output with synchronised audio, while none of the major labs had yet made a comparable model broadly available. Movie Gen’s audio component was also notable for tackling a problem — sound generated in sync with generated video, rather than added afterward — that most contemporaneous video models did not attempt.