Read the full write-up
Notion blog post with full details
Results
Approach
DeepScaleR iteratively scales Deepseek’s GRPO algorithm from 8K to 16K to 24K context length for thinking, trained on top of DeepSeek-R1-Distill-1.5B on math competition problems. See thecookbooks/math cookbook (single-turn math with \boxed{} answers) for the AgentFlow-based reproducer.
Released: February 2025
