Rabbit Holes a commonplace book,kept in the open
Tag · 1 folios

DeepSeek

A pencil mark from the long tail. Browse the hand-chosen subjects for broader paths through the book.

443 Link

There is still Alpha to be had in good training data.

source huggingface.co

Post-Training PipelineIn this release, we refrain from introducing novel post-training algorithms. The overall recipe follows the standard paradigm of supervised fine-tuning (SFT) followed by reinforcement learning (RL) and on-policy distillation (OPD) [Gu et al., 2024; Lu and Lall, 2025], without algorithmic modifications beyond well-established practices. Instead, our efforts are concentrated almost entirely on what the…

keep reading — 1 min ↗
AIDeepSeek 4 threads →