Guides
What to decide before you spend training compute, written from the runs on this site. Every number links to the curve it came from.
- 2026-09-27How to cut model training cost in half
Where training compute actually goes, why the data decision is the one lever that halves it, and what we measured: up to 63% less compute for the same held-out loss, with the receipts.
- 2026-09-27When curating training data backfires
The results nobody publishes: targets where selecting data toward the goal made the model worse, a frontier method that used three times the compute to lose to random data, and how to know before you train.
- 2026-09-27Data selection for pretraining, done so you can check it
The procedure behind every number on this site: samples, a verdict, a plain importance-weighting selector, a keep list computed where the data lives, and a two-arm test that produces a receipt.
- 2026-09-27Training cost calculator
What a training run costs in GPUs, hours and dollars, and what a measured reduction in compute is worth on your cluster. Projection, labelled as one.