Model
1 entry · February 2024
The 7B model reached 51.7% on the MATH benchmark without external tools, and its GRPO training method later underpinned DeepSeek-R1's reasoning training.
Ideas & essays · Models & capabilities