DeepMind's Veo 2 launches a week after OpenAI's Sora, claiming 4K output
Veo 2 claims 4K resolution and multi-minute generation, but the public VideoFX waitlist it launched into capped output at 720p and eight seconds.
- Models & capabilities
- Notable
Google DeepMind introduced Veo 2, an updated version of the text-to-video model it had unveiled in May 2024, a week after OpenAI made Sora publicly available. Google said Veo 2 could generate video at resolutions up to 4K and extending to several minutes in length, with improved understanding of real-world physics and human movement, and finer control over cinematography — letting users specify lens choice, camera movement and effects such as shallow depth of field in a prompt. In blind evaluations by human raters against competing video models, Google said Veo 2 came out ahead, though it did not name the models it was compared against or publish the evaluation methodology.
As with the original Veo, the capabilities Google demonstrated were well ahead of what was actually available to try. Veo 2 launched through an expanded waitlist for VideoFX, Google Labs’ experimental video tool, where public access was capped at 720p resolution and eight-second clips — far short of the 4K, multi-minute output shown in Google’s own demonstration videos. Google said broader access, including integration into YouTube Shorts, would follow in 2025. The same announcement introduced an updated Imagen 3 image model and Whisk, a new Labs tool that combined Imagen with Gemini’s visual understanding to let users remix and recombine images.
The release read as a direct, fast-follow response to Sora’s public launch rather than an independently timed announcement, continuing a pattern in which Google’s headline demonstrations ran well ahead of its shipping products — a gap that had also marked the original Veo’s release seven months earlier.