Google DeepMind shows Genie 2, an image-to-playable-3D-world model
Diffusion model turns a single prompt image into an explorable, physics-consistent 3D environment for training AI agents, kept to a research preview.
- Models & capabilities
- Notable
Google DeepMind showed Genie 2, a model that turns a single image or text-generated image into a playable, three-dimensional environment a person or an AI agent can navigate with keyboard and mouse input. Where its predecessor Genie, released earlier in 2024, generated flat 2D game-like worlds, Genie 2 produced explorable 3D scenes and was pitched not as a game engine but as a tool for training and evaluating other AI agents.
DeepMind said the model could hold a generated world consistent for up to a minute of continuous interaction, rendering it from multiple camera perspectives — first-person, third-person, isometric — and modelling physics, character animation, lighting and object interactions along the way. It could also generate several different continuations from the same starting frame, producing counterfactual versions of a scene, and demonstrated some memory of objects that had moved off-screen and back into view. DeepMind said an early version of its SIMA generalist game-playing agent could already act inside Genie 2’s generated worlds, framing the model as infrastructure for training embodied agents cheaply, in synthetic environments that did not require hand-built game levels.
DeepMind was explicit that the work was early: it said both the environments themselves and the agents acting within them still had “substantial room for improvement,” and that visual consistency and detail degraded over longer interactions. As with several of Google’s late-2024 model demonstrations, Genie 2 was not released as a public product; access was limited to a small group of researchers, with DeepMind describing it as a research preview rather than a shipping tool.
The release fit a broader pattern in frontier labs’ late-2024 output — interactive “world models” as a distinct research direction from both language models and video generators, aimed less at consumer output than at giving agents cheap simulated environments to train and be tested in.