443 Link
12 Sept 2026
There is still Alpha to be had in good training data.
source huggingface.cokeep reading — 1 min ↗Post-Training PipelineIn this release, we refrain from introducing novel post-training algorithms. The overall recipe follows the standard paradigm of supervised fine-tuning (SFT) followed by reinforcement learning (RL) and on-policy distillation (OPD) [Gu et al., 2024; Lu and Lall, 2025], without algorithmic modifications beyond well-established practices. Instead, our efforts are concentrated almost entirely on what the…