Benchmarks
Half the compute for the same model. Weeks back on every run. Measured before the money was spent.
saved per training run on a 100,000-GPU cluster
and about six weeks back on the calendar for every run
Every 1% of compute on a run this size is worth about $4M. Half of it is worth the race.
Projection: $2 per GPU-hour, a 90-day run, and our measured reduction of about half carried to this scale. Measured so far on models up to 41.7M parameters; a pilot on your own job verifies it on yours.
19 verdicts. Every one finished on the side we named. 15 were sealed and timestamped before their first training step.
More than 300 full training runs sit behind these lines, four seeds per arm. The loss column is the curated run's best held-out loss minus the random baseline's; below zero, curating helped.
Sealed: the call was fixed and timestamped before the first run, then checked against the outcome. Measured: the first four cells, where the rule was found and its boundary located; no prior call existed, so they are listed and not counted as sealed. Receipt: a single-seed run through the public service, kept as a check on the service, not as a battery result.
Same quality, about half the training
Training compute the curated run needed to reach the best quality the random-data baseline ever reached.
Training steps the curated run needs to reach the best held-out quality the random-data baseline ever achieved, averaged over four runs per domain. On the DON'T job the curated model never reaches the baseline at any amount of training, as the verdict said. A near-control domain (9 to 22% less) is left out.
At equal compute, the right call is worth up to a third
Quality gained by following the call, against training on a random sample of the same size.
Quality is perplexity on held-out target text, the standard accuracy measure for language models. On the DON'T job, curating would have cost 8%; the verdict kept the baseline. One near-control domain (2% better, GO) is left out. Every call was made before training.
Against the frontier
Two targets where our call was GO. The most advanced published curation methods, run at full strength, against our call followed by a plain published selector. Same model, same training budget, four seeds each.
Quality is perplexity on held-out target text against the random-data baseline at the same training budget. MATES re-probes the model during training; its runs took about 1,900 GPU-seconds against about 480 for every other arm on the same hardware, so its compute is roughly three times the others. Its influence predictor was verified as working in every round, so this is the method at its best, not a weakened copy. Forty runs, preregistered, every comparison 4 of 4 seeds with the interval excluding zero.
Every number above, with the training curve it came from.
Hover a curve to read it. The gray line is the model trained on a random sample; the blue line is the same model on the curated sample. The dot is where the curated run first reached the best the baseline ever did.
Eight domains of real text, from scratch
Every domain: four seeds per arm, six thousand steps, the same model and budget on both sides. The only difference is which quarter of the pool the model trained on.
Medical (PubMed)
- Headline
- 63% less compute
- Reached the baseline's best
- step 2,214 of 5,999
- Best held-out loss
- −0.194 lower than baseline · 95% [−0.208, −0.175]
- Quality, perplexity
- 17.6% better
- Per seed
- −0.165 −0.194 −0.204 −0.212
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
Legal
- Headline
- 60% less compute
- Reached the baseline's best
- step 2,176 of 5,400
- Best held-out loss
- −0.190 lower than baseline · 95% [−0.208, −0.169]
- Quality, perplexity
- 17.3% better
- Per seed
- −0.186 −0.212 −0.204 −0.158
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
Recipes
- Headline
- 58% less compute
- Reached the baseline's best
- step 2,517 of 5,999
- Best held-out loss
- −0.191 lower than baseline · 95% [−0.201, −0.181]
- Quality, perplexity
- 17.4% better
- Per seed
- −0.206 −0.195 −0.187 −0.176
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
Medical, tiny sample
- Headline
- 58% less compute
- Reached the baseline's best
- step 2,523 of 5,999
- Best held-out loss
- −0.170 lower than baseline · 95% [−0.190, −0.150]
- Quality, perplexity
- 15.6% better
- Per seed
- −0.152 −0.192 −0.148 −0.188
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
Python documentation
- Headline
- 52% less compute
- Reached the baseline's best
- step 2,681 of 5,600
- Best held-out loss
- −0.420 lower than baseline · 95% [−0.520, −0.325]
- Quality, perplexity
- 34.3% better
- Per seed
- −0.333 −0.316 −0.582 −0.448
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
News
- Headline
- 35% less compute
- Reached the baseline's best
- step 3,123 of 4,800
- Best held-out loss
- −0.071 lower than baseline · 95% [−0.090, −0.052]
- Quality, perplexity
- 6.9% better
- Per seed
- −0.051 −0.096 −0.085 −0.053
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
Encyclopedia (near-control)
- Headline
- 9% less compute
- Reached the baseline's best
- step 5,481 of 5,999
- Best held-out loss
- −0.023 lower than baseline · 95% [−0.032, −0.011]
- Quality, perplexity
- 2.3% better
- Per seed
- −0.026 −0.005 −0.028 −0.034
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
Encyclopedia, tiny sample
- Headline
- never catches up
- Best held-out loss
- +0.076 higher than baseline · 95% [+0.069, +0.084]
- Quality, perplexity
- 7.9% worse if curated
- Per seed
- +0.083 +0.084 +0.071 +0.066
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
Loss is cross-entropy on held-out target text, evaluated every 200 steps. Compute saved is where the curated run's mean curve first reaches the best loss the baseline's mean curve ever reached, against the step where the baseline reached it. The 95% interval is a 20,000-draw bootstrap over the four paired seeds. Every call was fixed and timestamped before the first run.
The same calls, on a model eight times larger
Three of the domains rerun on a model with eight times the non-embedding capacity. Same pool, same selections, same frozen calls. Nothing retuned.
Medical (PubMed)
- Headline
- 56% less compute
- Reached the baseline's best
- step 2,180 of 5,000
- Best held-out loss
- −0.210 lower than baseline · 95% [−0.230, −0.191]
- Quality, perplexity
- 18.9% better
- Per seed
- −0.239 −0.184 −0.213 −0.203
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
Encyclopedia (near-control)
- Headline
- 22% less compute
- Reached the baseline's best
- step 4,690 of 5,999
- Best held-out loss
- −0.023 lower than baseline · 95% [−0.040, −0.011]
- Quality, perplexity
- 2.3% better
- Per seed
- −0.047 −0.007 −0.018 −0.019
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
Encyclopedia, tiny sample
- Headline
- never catches up
- Best held-out loss
- +0.063 higher than baseline · 95% [+0.038, +0.077]
- Quality, perplexity
- 6.4% worse if curated
- Per seed
- +0.026 +0.078 +0.069 +0.077
- Runs
- 4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept
41.7M parameters, four seeds per arm, six thousand steps. The medical gain grew, the near-control gain came out identical to the small model, and the DON'T stayed a DON'T.
Image and text together
A 256M-parameter vision-language model, continued on curated image-text pairs. The first completed target of a 48-run battery, and the same target run end to end through the public service.
Charts and plots
- Headline
- 67% less compute
- Reached the baseline's best
- step 2,696 of 8,249
- Best held-out loss
- −0.057 lower than baseline
- Quality, perplexity
- 5.5% better
- Per seed
- −0.057
- Runs
- 1 seed × 2 arms · 8,261 steps · 35 evaluations · 66,087 documents kept
Charts and plots, through the public API
- Headline
- 65% less compute
- Reached the baseline's best
- step 2,922 of 8,260
- Best held-out loss
- −0.057 lower than baseline
- Quality, perplexity
- 5.5% better
- Per seed
- −0.057
- Runs
- 1 seed × 2 arms · 8,261 steps · 35 evaluations · 66,087 documents kept
Single seed so far; these are signs, not intervals. Through the service, the verdict was issued and sealed before training, the selection list came from the API, and the run happened on a different GPU stack and framework; its gap matched our own run of the same target to three decimals. The remaining 42 runs of the battery follow.
Every verdict was issued and timestamped before its training run began. Models up to 41.7M parameters on real text, and a 256M-parameter vision-language model in progress; larger models are the next sealed tests. Every run, including the ones we lost, sits in an audit ledger available under NDA.