Generative media
The rise of AI image, video and audio generation — from DALL·E and Stable Diffusion to Sora and Veo 3 — and the misuse, from nonconsensual deepfakes to copyright disputes, that tracked its quality.
Generative media is the thread of models that make images, video and audio rather than text. It opened with OpenAI’s DALL·E and CLIP, the pairing that connected language to pictures, and accelerated through DALL·E 2, Midjourney’s Discord beta and the open release of Stable Diffusion, which put photorealistic image generation in anyone’s hands.
Video was the next frontier. OpenAI previewed Sora in early 2024 and released it publicly that December; Google answered with Veo 3, and OpenAI’s native image generation in GPT-4o set off a viral Studio Ghibli-style wave. By late 2025 Sora 2 had turned generation into a social feed.
Quality brought misuse in step. The same tools were used to make nonconsensual sexual deepfakes, and xAI’s image features were turned to mass-produced abuse imagery — harms that put generative media at the centre of the deepfake-law and consent debates traced in other threads. The tension is constant across it: each gain in fidelity widened both the creative use and the capacity for harm.
OpenAI releases Jukebox
Compressed raw audio into discrete codes with a multi-scale VQ-VAE, then modelled those with autoregressive transformers to keep coherence over several minutes.
Models & capabilities
OpenAI announces DALL·E and CLIP
Two multimodal models released on the same day: one that generated images from text captions, one that learned to match images to language.
Models & capabilities · Open weights & ecosystem
A deepfake Tom Cruise goes viral on TikTok
Belgian VFX artist Chris Ume paired face-swap software with impersonator Miles Fisher's performance; the clips passed 11 million TikTok views within days.
Culture & impact · Security & misuse
LAION releases LAION-5B image-text dataset
LAION released LAION-5B, then the largest freely available image-text dataset (5.85bn pairs), which became the training corpus for Stable Diffusion and other open image models.
Open weights & ecosystem
OpenAI releases DALL·E 2
Higher resolution and photorealism than the original DALL·E, released first to roughly 400 trusted users pending a public waitlist.
Models & capabilities
Google unveils Imagen text-to-image diffusion model
Google Research reported higher photorealism scores than DALL·E 2 in side-by-side human evaluation, but declined to release code or a public demo.
Models & capabilities
Midjourney opens its beta on Discord
Image requests and results were posted in public chat channels by default, turning generation into a shared, watchable activity rather than a private tool.
Models & capabilities · Culture & impact
Stable Diffusion is released to the public
Weights published under a permissive licence and runnable on a consumer graphics card, with community fine-tunes and graphical front-ends appearing within weeks.
Open weights & ecosystem · Models & capabilities
An AI image wins the Colorado State Fair art prize
Jason Allen's Midjourney piece 'Théâtre D'opéra Spatial' took first place in the fair's digital arts category and a $300 prize, and was later denied US copyright registration.
Culture & impact
OpenAI open-sources Whisper
Trained on 680,000 hours of web-scraped audio and released under the MIT licence, an unusual openness for a company otherwise moving toward closed models.
Open weights & ecosystem · Models & capabilities
Meta and Google race to text-to-video
Make-A-Video and Imagen Video appeared within a week of each other, both as research previews with no public access.
Models & capabilities
DeepMind introduces MusicLM, a text-to-music generation model
The model composed several minutes of coherent audio from a written prompt or a hummed melody, released as a research paper rather than a product.
Models & capabilities
Midjourney releases V5
The update was widely noted for fixing AI image generation's most-mocked failure, malformed hands, and for images some viewers mistook for photographs.
Models & capabilities
Meta releases Segment Anything Model (SAM)
Trained on 1.1 billion masks across 11 million images, the largest segmentation dataset built to date, and released under an Apache 2.0 licence.
Open weights & ecosystem · Models & capabilities
'Heart on My Sleeve' AI Drake/Weeknd song goes viral, then pulled from streaming
Released 4 April 2023 by the anonymous TikTok account Ghostwriter977, the track reached about 600,000 Spotify streams and 15 million TikTok views before Universal Music Group had it pulled.
Culture & impact
Runway launches Gen-2 text-to-video to the public
Clips were capped at four seconds and reviewers called the output 'more a novelty or toy than a genuinely useful tool,' citing melting objects and low framerates.
Models & capabilities
Synthesia raises $90M Series C, becomes AI-video unicorn
The London-based company said it had over 50,000 customers and had generated more than 15 million videos, up from a $300M valuation 18 months earlier.
Money & business
ElevenLabs raises $19M Series A
The round valued the company at $99M; alongside it, ElevenLabs launched a free tool letting anyone check whether audio was generated with its own voice models.
Money & business · Models & capabilities
Stability AI releases Stable Diffusion XL
The two-stage model, combining a 3.5B-parameter base and 6.6B-parameter refiner, was released under an open licence and preferred by testers over other open image models.
Models & capabilities · Open weights & ecosystem
Meta releases AudioCraft (MusicGen, AudioGen, EnCodec)
Meta published weights and code for all three models, extending its open-release strategy from language models into audio and music generation.
Open weights & ecosystem · Models & capabilities
OpenAI releases DALL·E 3
ChatGPT itself was positioned as the prompt engineer, drafting detailed image prompts from short user requests before rendering them; rollout began the following month.
Models & capabilities
Pika Labs launches Pika 1.0, raises $55M
Founded by two Stanford computer-science PhD students, the video-editing app had already drawn some 500,000 users through its earlier Discord-only release.
Models & capabilities · Money & business
Suno launches public web app
Cambridge, Massachusetts-based Suno, founded by ex-Kensho engineers, opens its AI song-generation web app to the public, alongside a Microsoft Copilot integration.
Models & capabilities
Sexual deepfakes of Taylor Swift spread on X
One post was viewed over 47 million times before removal; X temporarily blocked searches of her name and the White House urged Congress to legislate.
Security & misuse · Culture & impact
OpenAI previews Sora, a text-to-video model
Clips up to a minute long held their subjects consistent across camera moves; OpenAI showed samples but gave no public access for ten months.
Models & capabilities · Culture & impact
Suno releases v3
The company called it its first model able to produce radio-quality output, letting free users generate two-minute songs and paying users four-minute tracks.
Models & capabilities
Udio launches in beta
Backed by Andreessen Horowitz and artists including will.i.am, the free text-to-music beta arrived weeks after rival Suno's v3, intensifying competition in AI music.
Models & capabilities
Microsoft Research shows VASA-1, real-time talking-face generation
Trained to render 512x512 video at up to 40 frames per second from one photo and an audio clip; Microsoft said it had no plans to release a demo or model.
Models & capabilities · Security & misuse
OpenAI launches GPT-4o with real-time voice
A single model handling text, vision and audio end to end, with conversational latency — and a voice that led to a public dispute with Scarlett Johansson.
Models & capabilities · Culture & impact
Google DeepMind unveils Veo, a text-to-video generation model
Generates 1080p video over a minute long from text and image prompts, watermarked with SynthID; launched in private preview via waitlist rather than public release.
Models & capabilities
Suno raises $125M Series B at $500M valuation
Suno said 10 million people had made music on the platform in its first eight months, ahead of a round led by Lightspeed with backing from Nat Friedman and Daniel Gross.
Money & business
Luma AI launches Dream Machine
Free-to-use video model generating five-second clips at 1360x752 from text or image prompts, drawing early comparisons to OpenAI's unreleased Sora.
Models & capabilities
Stability AI ships Stable Diffusion 3 Medium weights
The 2-billion-parameter weights ran on consumer GPUs but were licensed non-commercially, with a separate paid Enterprise tier for large-scale business use.
Models & capabilities · Open weights & ecosystem
Runway unveils Gen-3 Alpha
Trained jointly on video and images on new infrastructure, it added fine-grained temporal control and C2PA provenance labelling absent from Runway's Gen-2 model.
Models & capabilities
Record labels sue Suno and Udio
Suno later conceded training on copyrighted recordings but argued the use was transformative fair use, comparable to a person learning to write by reading.
Courts & copyright
Black Forest Labs releases FLUX
The 12-billion-parameter FLUX.1 suite shipped in three tiers — a paid API model, an open non-commercial model and an Apache-licensed fast model — funded by $31 million in seed money.
Open weights & ecosystem · Models & capabilities
Google opens Imagen 3 access to all US users
Imagen 3, previewed at Google I/O in May, becomes available through ImageFX to all US users, with improved text rendering and fewer visual artefacts.
Models & capabilities
NotebookLM adds Audio Overviews, AI-generated podcast-style summaries
Two AI hosts discuss a user's uploaded documents in an unscripted-sounding conversation; Google called it experimental and English-only at launch.
Models & capabilities
Meta releases Movie Gen video and audio foundation models
A 30B-parameter video model paired with a 13B-parameter audio model produced clips up to 16 seconds with synchronised sound; Meta called it research, not a product.
Models & capabilities
Stability AI releases Stable Diffusion 3.5
Stability AI releases Stable Diffusion 3.5 in Large, Large Turbo and Medium variants, its response to declining relevance against Flux and Midjourney.
Models & capabilities · Open weights & ecosystem
Coca-Cola's AI-generated Christmas ad draws backlash over 'soulless' imagery
Coca-Cola's AI-generated holiday ad, remaking its classic 1995 truck commercial, was widely mocked online as uncanny and 'soulless', becoming a flashpoint in advertising's AI debate.
Culture & impact
Tencent releases HunyuanVideo
At 13 billion parameters, Tencent called it the largest open-weight video model, and said blind evaluators rated its motion quality above Runway Gen-3 and Luma 1.6.
Open weights & ecosystem · Models & capabilities
OpenAI releases Sora publicly
Every clip carried a visible watermark and embedded C2PA provenance metadata; access launched in the US and Canada only, excluding the UK and EU.
Models & capabilities
OpenAI publishes the Sora system card
Published alongside Sora's public launch, it disclosed testing by red-teamers in nine countries on more than 15,000 generations, and that the model still struggles with realistic physics.
Safety & alignment · Models & capabilities
DeepMind's Veo 2 launches a week after OpenAI's Sora, claiming 4K output
Veo 2 claims 4K resolution and multi-minute generation, but the public VideoFX waitlist it launched into capped output at 720p and eight seconds.
Models & capabilities
Synthesia raises $180M Series D, valuation doubles to $2.1B
The round, led by NEA with NVentures among existing backers, valued the AI-avatar video firm at roughly double its 2023 Nvidia-backed Series C mark.
Money & business
ChatGPT's image generator sets off a Ghibli wave
Native image generation produced a flood of Studio Ghibli pastiche, adding a million users in an hour and reopening the style-copyright argument.
Culture & impact · Models & capabilities · Courts & copyright
OpenAI adds native image generation to GPT-4o
Images were generated natively by GPT-4o's own architecture rather than by a separate diffusion model, improving text rendering and editing of existing images.
Models & capabilities
Runway releases Gen-4
The model generates the same character, object or location across separate video shots from a single reference image, addressing consistency limits that had confined earlier AI video to short isolated clips.
Models & capabilities
Google I/O puts Gemini into search and ships Veo 3
AI Mode rolled out to all US Search users, and Veo 3 became the first widely-used video model to generate synchronised dialogue and sound effects alongside the picture.
Models & capabilities
ByteDance launches Seedance 1.0 video generation model
ByteDance said the text- and image-to-video model topped Artificial Analysis's leaderboards and generated a five-second 1080p clip in about 41 seconds.
Models & capabilities
Disney and Universal sue Midjourney
The 110-page complaint sought $150,000 per infringed work over characters including Darth Vader, Elsa and the Minions, which Midjourney's generator would reproduce on request.
Courts & copyright
xAI's Grok Imagine 'spicy mode' used to make nonconsensual Taylor Swift deepfakes
A Verge reporter got explicit video of the singer on a first, unjailbroken attempt, unlike rival tools from Google and OpenAI which blocked celebrity nudity outright.
Culture & impact · Security & misuse · Safety & alignment
Google releases Gemini 2.5 Flash Image, nicknamed 'Nano Banana'
Priced at $0.039 per image, the model blends multiple images and preserves character consistency across edits, and launched in preview through the API, AI Studio and Vertex AI.
Models & capabilities
Tencent open-sources Hunyuan Image 3.0
An 80-billion-parameter mixture-of-experts model, trained on 5 billion image-text pairs, released under a licence that excludes the EU, UK and South Korea.
Open weights & ecosystem · Models & capabilities
OpenAI publishes Sora 2 system card
The document, released alongside the app, set out cameo consent controls, C2PA provenance metadata and likeness-detection safeguards without disclosing training data or red-team pass rates.
Safety & alignment
Sora 2 launches as a social video app
An invite-only iOS feed of short AI videos with a consent-based cameo feature for real likenesses, launched with an opt-out copyright policy OpenAI reversed within days.
Models & capabilities · Culture & impact
MiniMax releases Hailuo 2.3 video model
MiniMax kept pricing level with the prior Hailuo 02 model while improving character movement, facial micro-expressions and stylised rendering.
Models & capabilities
Disney and OpenAI reach a Sora content agreement
Disney agreed to license over 200 characters from Marvel, Pixar and Star Wars for one-year exclusive use in Sora and to invest $1bn in OpenAI, without covering talent likenesses or voices.
Courts & copyright · Money & business
ByteDance launches Seedance 1.5 Pro with joint audio-video generation
The model generates video and audio through a single diffusion transformer rather than a separate pass, aiming for accurate lip-sync across eight languages.
Models & capabilities
OpenAI updates ChatGPT image generation to GPT Image 1.5
OpenAI moved the launch up from a planned January date, part of a competitive scramble after Google's Gemini 3 and its 'Nano Banana Pro' image tool led leaderboards.
Models & capabilities
xAI's Grok Edit Image feature used to mass-produce nonconsensual sexualized images
The Center for Countering Digital Hate estimated over 3 million sexualized images were generated in 11 days, roughly 23,000 depicting apparent children, before X restricted the feature to paid users.
Security & misuse · Culture & impact
UK Ofcom opens formal investigation into X over Grok deepfakes
Ofcom cited possible failures on illegal content, risk assessment and child safety duties, warning of fines up to £18 million or 10% of global revenue.
Security & misuse · Government & policy
California AG issues cease-and-desist to xAI over Grok deepfakes
Attorney General Rob Bonta invoked the state's new civil deepfake-pornography statute and CSAM law, giving xAI five days to confirm it had stopped the conduct.
Security & misuse · Courts & copyright
Google DeepMind releases Genie 3 world model
Genie 3, previously a limited research preview, opened to Google AI Ultra subscribers in the US as Project Genie, generating explorable 3D worlds from a text or image prompt.
Models & capabilities
OpenAI shuts down Sora video app after deepfake and 'AI slop' backlash
OpenAI said it would wind the app down roughly six months after its September 2025 launch, pulling it from app stores for new users while the full web and API shutdown followed months later.
Culture & impact · Safety & alignment
OpenAI discontinues the Sora app and web product
The app and web product closed after roughly seven months; OpenAI said it was reallocating computing resources toward coding tools and enterprise products, with the API to follow in September.
Culture & impact · Models & capabilities
Google introduces Gemini Omni, a unified any-input-to-video model
Gemini Omni Flash rolled out free inside YouTube Shorts Remix, letting users restyle or insert themselves into existing videos, with every output carrying a SynthID watermark.
Models & capabilities
OpenAI adds SynthID watermarking and a public content-verification tool
The invisible watermark, developed by a rival lab, is paired with C2PA metadata credentials because OpenAI said neither signal survives image manipulation on its own.
Culture & impact
ByteDance unveils Seedance 2.5, generating native 30-second video clips
The model generates single clips up to 30 seconds long, including scene and tempo changes, without stitching together separate shots as earlier video models required.
Models & capabilities · Benchmarks & progress
MiniMax unveils Hailuo 3.0 (H3) video model with native 2K and synced audio
The model generates synchronised dialogue, sound effects and ambient audio alongside native 2K video in a single pass, and accepts up to nine reference images for consistency.
Models & capabilities
Deezer says AI-generated tracks exceed half of daily music uploads
The share climbed from 10% of roughly 10,000 daily uploads in January 2025 to over half of about 90,000 daily uploads by June 2026, Deezer's own detection tool found.
Culture & impact