◷ Incident Log · radical candor, in public
7 EventsNo Western lab publishes this mid-run. Restarts, OOMs, broken data pipelines — and the shutdown fingerprint at step 30. Read the full detective story in chapter 08, and see the restart artifacts in the KL chart.
✓ Capstone · read any RL dashboard like an engineer
Normal vs alarm, pocket edition
Normal: avg@n + reward + benchmarks climbing; KL gently rising between restart resets; entropy drifting; step time growing with context; infra_error near zero. Alarm: reward/pass divergence; grad_norm spike with no notice; infra_error spike masquerading as a capability dip; passrate/one swallowing the batch with no curriculum refresh.
| Figure | Pro (step 30) | Flash (step 30) |
|---|---|---|
| avg@n (Δ vs step 1) | 0.633 (▲0.068) | 0.644 (▲0.130) |
| Cost so far | $2,620,670 | $854,044 |
| Tokens (step / total) | 3.43B / 75B | 3.7B / 81.4B |
| Context / turns | ~137k / 43.6 | ~148k / 68.5 |
| Step time | 6h26m | 3h27m |
| DeepSWE / in-house / automation | 72.57 / 65.43 / 53.10 | 65.68 / 62.87 / 52.70 |
Sources: live dashboard snapshot captured 2026-09-21 (mimo.xiaomi.com/rl); final-step values from the run's step-30 state. Inferred definitions are labeled as inference in the text.