Stability AI previews Stable Diffusion 3
The suite spans 800 million to 8 billion parameters and combines a diffusion transformer with flow matching, aimed chiefly at fixing text rendering and multi-subject prompts.
- Models & capabilities
- Minor
Stability AI previewed Stable Diffusion 3, moving its flagship text-to-image line to a diffusion transformer architecture combined with a “flow matching” training method, in place of the U-Net-based diffusion approach used in earlier Stable Diffusion versions. The company opened a waitlist for early access rather than releasing weights immediately, saying the preview period was intended to gather feedback on performance and safety ahead of a wider release.
The suite spanned a range of model sizes from 800 million to 8 billion parameters, allowing users to trade capability for lower compute requirements. Stability said the architecture change targeted specific, longstanding weaknesses of diffusion image models: handling prompts describing multiple distinct subjects, and rendering legible text within generated images, both areas where earlier Stable Diffusion versions and rivals had struggled.
The announcement came as Stability AI was working through a period of internal instability, including senior departures and reported financial strain, making a credible architectural advance on its core product commercially significant for the company as well as technically notable. Public weights for Stable Diffusion 3 did not follow immediately from this preview; a wider release of model weights came several months later, after further safety testing and adjustments prompted partly by problems users found with anatomical accuracy in the initial preview outputs.