Timeline

Meta and Google race to text-to-video

Make-A-Video and Imagen Video appeared within a week of each other, both as research previews with no public access.

  • Models & capabilities
  • Minor

Meta AI announced Make-A-Video, a system that generated short video clips — up to five seconds, without sound — from a text prompt, extending the text-to-image approach that had produced DALL·E 2 and Stable Diffusion earlier that year into a new modality. The model learned what objects and scenes look like from paired text-image data, and how the world moves from unlabelled video footage, combining the two without needing large volumes of text-labelled video, which barely existed at the scale needed for training. Zuckerberg’s own demonstration clips — a teddy bear painting a self-portrait, a spacecraft landing on Mars — set the tone for how the technology was received: proof of concept, visibly imperfect, obviously heading somewhere.

Six days later, Google detailed Imagen Video, its own text-conditioned video diffusion model, built on the Imagen text-to-image architecture it had introduced earlier in 2022. Google’s system generated higher-resolution output — up to 1280×768 at 24 frames per second — and demonstrated a range of artistic styles, 3D consistency and legible on-screen text, capabilities Meta’s system had not shown.

Neither system was released for public use; both were research previews documented in blog posts and papers, with no API, no waitlist and no weights released. The near-simultaneous announcements made clear that, having largely converged on text-to-image within the same year, the frontier labs had already moved their competitive attention to video — a modality that took substantially longer to reach usable public products than image generation had, with OpenAI’s Sora and comparable systems from other labs not arriving until well over a year later.