StepFun releases Step 3.5 Flash, topping several reasoning benchmarks
The 196-billion-parameter model, of which only about 11 billion activate per token, was released under an Apache 2.0 licence and scored 97.3% on AIME 2025.
- Open weights & ecosystem
- Models & capabilities
- Benchmarks & progress
- Minor
Chinese lab StepFun released Step 3.5 Flash, a sparse mixture-of-experts model with 196 billion total parameters of which only about 11 billion activate for any given token, under an Apache 2.0 open-weights licence. On StepFun’s reported figures the model scored 97.3% on AIME 2025 and 85.4% on IMOAnswerBench, alongside 74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0, leading several reasoning and agentic benchmarks against larger rivals.
The comparison that drew attention was scale: StepFun’s model beat results reported for substantially bigger Chinese open-weight releases, including Moonshot AI’s roughly 1-trillion-parameter Kimi K2.5 and DeepSeek’s 671-billion-parameter V3.2, as well as models from Zhipu AI and MiniMax, on several of these benchmarks. StepFun said the design prioritised strong reasoning ability within an efficient context window and fast inference speed rather than raw parameter count.
The release landed amid intensifying competition among Chinese labs releasing openly licensed models around the Lunar New Year period, following DeepSeek’s earlier releases that had already reset expectations about how much capability a comparatively compact, open-weight model could deliver against larger closed systems from Western labs.