Echo Program

Benchmarks

Half the compute for the same model. Weeks back on every run. Measured before the money was spent.

$216M

saved per training run on a 100,000-GPU cluster

and about six weeks back on the calendar for every run

Cluster
90-day run
Saved
Earlier
1,000 GPUs
$4.3M
$2.2M
6 weeks
10,000 GPUs
$43M
$22M
6 weeks
100,000 GPUs
$432M
$216M
6 weeks

Every 1% of compute on a run this size is worth about $4M. Half of it is worth the race.

Projection: $2 per GPU-hour, a 90-day run, and our measured reduction of about half carried to this scale. Measured so far on models up to 41.7M parameters; a pilot on your own job verifies it on yours.

The ledger

19 verdicts. Every one finished on the side we named. 15 were sealed and timestamped before their first training step.

More than 300 full training runs sit behind these lines, four seeds per arm. The loss column is the curated run's best held-out loss minus the random baseline's; below zero, curating helped.

2026-08-30
Synthetic shift, weak
DON'T · +0.022
random won
measured
2026-08-30
Synthetic shift, weak
DON'T · +0.015
random won
measured
2026-08-30
Synthetic shift, strong
GO · −0.085
curated won
measured
2026-08-30
Synthetic shift, strong
GO · −0.100
curated won
measured
2026-08-30
Synthetic shift, weak, fresh draw
DON'T · +0.020
random won
sealed
2026-08-30
Synthetic shift, weak, fresh draw
DON'T · +0.018
random won
sealed
2026-08-30
Synthetic shift, strong, fresh draw
GO · −0.064
curated won
sealed
2026-08-30
Synthetic shift, strong, fresh draw
GO · −0.117
curated won
sealed
2026-09-05
Medical (PubMed)
GO · −0.194
curated won
sealed
2026-09-05
Python documentation
GO · −0.420
curated won
sealed
2026-09-05
Legal
GO · −0.190
curated won
sealed
2026-09-05
News
GO · −0.071
curated won
sealed
2026-09-05
Recipes
GO · −0.191
curated won
sealed
2026-09-05
Encyclopedia (near-control)
GO · −0.023
curated won
sealed
2026-09-05
Encyclopedia, tiny sample
DON'T · +0.076
random won
sealed
2026-09-05
Medical, tiny sample
GO · −0.170
curated won
sealed
2026-09-05
Medical (PubMed)
GO · −0.210
curated won
sealed
2026-09-05
Encyclopedia (near-control)
GO · −0.023
curated won
sealed
2026-09-05
Encyclopedia, tiny sample
DON'T · +0.062
random won
sealed
2026-09-24
Charts and plots, via the public API
GO · −0.057
curated won
receipt

Sealed: the call was fixed and timestamped before the first run, then checked against the outcome. Measured: the first four cells, where the rule was found and its boundary located; no prior call existed, so they are listed and not counted as sealed. Receipt: a single-seed run through the public service, kept as a check on the service, not as a battery result.

2026-09-05

Same quality, about half the training

Training compute the curated run needed to reach the best quality the random-data baseline ever reached.

Target domain
Our call
Outcome
Medical (PubMed)
GO
63% less
Legal
GO
60% less
Recipes
GO
58% less
Medical, tiny sample
GO
58% less
Medical, 8× larger model
GO
56% less
Python documentation
GO
52% less
News
GO
35% less
Encyclopedia, tiny sample
DON'T
never catches up

Training steps the curated run needs to reach the best held-out quality the random-data baseline ever achieved, averaged over four runs per domain. On the DON'T job the curated model never reaches the baseline at any amount of training, as the verdict said. A near-control domain (9 to 22% less) is left out.

2026-09-05

At equal compute, the right call is worth up to a third

Quality gained by following the call, against training on a random sample of the same size.

Target domain
Our call
Outcome
Python documentation
GO
34% better
Medical (PubMed)
GO
18% better
Recipes
GO
17% better
Legal
GO
17% better
Medical, tiny sample
GO
16% better
News
GO
7% better
Encyclopedia, tiny sample
DON'T
8% loss avoided

Quality is perplexity on held-out target text, the standard accuracy measure for language models. On the DON'T job, curating would have cost 8%; the verdict kept the baseline. One near-control domain (2% better, GO) is left out. Every call was made before training.

2026-08-30

Against the frontier

Two targets where our call was GO. The most advanced published curation methods, run at full strength, against our call followed by a plain published selector. Same model, same training budget, four seeds each.

Synthetic shift, strong (A)
random baseline 4.604
MATES, full, with re-probing
≈3× compute
36.7% worse
BETR, classifier
1× compute
4.1% better
Ours: GO, then a plain selector
1× compute
8.5% better
Synthetic shift, strong (B)
random baseline 4.639
MATES, full, with re-probing
≈3× compute
28.2% worse
BETR, classifier
1× compute
5.1% better
Ours: GO, then a plain selector
1× compute
11.3% better

Quality is perplexity on held-out target text against the random-data baseline at the same training budget. MATES re-probes the model during training; its runs took about 1,900 GPU-seconds against about 480 for every other arm on the same hardware, so its compute is roughly three times the others. Its influence predictor was verified as working in every round, so this is the method at its best, not a weakened copy. Forty runs, preregistered, every comparison 4 of 4 seeds with the interval excluding zero.

The runs

Every number above, with the training curve it came from.

Hover a curve to read it. The gray line is the model trained on a random sample; the blue line is the same model on the curated sample. The dot is where the curated run first reached the best the baseline ever did.

2026-09-05

Eight domains of real text, from scratch

Every domain: four seeds per arm, six thousand steps, the same model and budget on both sides. The only difference is which quarter of the pool the model trained on.

Medical (PubMed)

GO
Small language model, trained from scratch
03,0006,0004.92baselinecurated
BaselineCuratedBaseline's best
Headline
63% less compute
Reached the baseline's best
step 2,214 of 5,999
Best held-out loss
−0.194 lower than baseline · 95% [−0.208, −0.175]
Quality, perplexity
17.6% better
Per seed
−0.165 −0.194 −0.204 −0.212
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

Legal

GO
Small language model, trained from scratch
03,0006,0005.11baselinecurated
BaselineCuratedBaseline's best
Headline
60% less compute
Reached the baseline's best
step 2,176 of 5,400
Best held-out loss
−0.190 lower than baseline · 95% [−0.208, −0.169]
Quality, perplexity
17.3% better
Per seed
−0.186 −0.212 −0.204 −0.158
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

Recipes

GO
Small language model, trained from scratch
03,0006,0005.18baselinecurated
BaselineCuratedBaseline's best
Headline
58% less compute
Reached the baseline's best
step 2,517 of 5,999
Best held-out loss
−0.191 lower than baseline · 95% [−0.201, −0.181]
Quality, perplexity
17.4% better
Per seed
−0.206 −0.195 −0.187 −0.176
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

Medical, tiny sample

GO
Small language model, trained from scratch
03,0006,0004.92baselinecurated
BaselineCuratedBaseline's best
Headline
58% less compute
Reached the baseline's best
step 2,523 of 5,999
Best held-out loss
−0.170 lower than baseline · 95% [−0.190, −0.150]
Quality, perplexity
15.6% better
Per seed
−0.152 −0.192 −0.148 −0.188
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

Python documentation

GO
Small language model, trained from scratch
03,0006,0004.45baselinecurated
BaselineCuratedBaseline's best
Headline
52% less compute
Reached the baseline's best
step 2,681 of 5,600
Best held-out loss
−0.420 lower than baseline · 95% [−0.520, −0.325]
Quality, perplexity
34.3% better
Per seed
−0.333 −0.316 −0.582 −0.448
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

News

GO
Small language model, trained from scratch
03,0006,0006.18baselinecurated
BaselineCuratedBaseline's best
Headline
35% less compute
Reached the baseline's best
step 3,123 of 4,800
Best held-out loss
−0.071 lower than baseline · 95% [−0.090, −0.052]
Quality, perplexity
6.9% better
Per seed
−0.051 −0.096 −0.085 −0.053
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

Encyclopedia (near-control)

GO
Small language model, trained from scratch
03,0006,0004.47baselinecurated
BaselineCuratedBaseline's best
Headline
9% less compute
Reached the baseline's best
step 5,481 of 5,999
Best held-out loss
−0.023 lower than baseline · 95% [−0.032, −0.011]
Quality, perplexity
2.3% better
Per seed
−0.026 −0.005 −0.028 −0.034
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

Encyclopedia, tiny sample

DON'T
Small language model, trained from scratch
03,0006,0004.47baselinecurated
BaselineCuratedBaseline's best
Headline
never catches up
Best held-out loss
+0.076 higher than baseline · 95% [+0.069, +0.084]
Quality, perplexity
7.9% worse if curated
Per seed
+0.083 +0.084 +0.071 +0.066
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

Loss is cross-entropy on held-out target text, evaluated every 200 steps. Compute saved is where the curated run's mean curve first reaches the best loss the baseline's mean curve ever reached, against the step where the baseline reached it. The 95% interval is a 20,000-draw bootstrap over the four paired seeds. Every call was fixed and timestamped before the first run.

2026-09-05

The same calls, on a model eight times larger

Three of the domains rerun on a model with eight times the non-embedding capacity. Same pool, same selections, same frozen calls. Nothing retuned.

Medical (PubMed)

GO
Language model with 8× the capacity, 41.7M parameters
03,0006,0004.82baselinecurated
BaselineCuratedBaseline's best
Headline
56% less compute
Reached the baseline's best
step 2,180 of 5,000
Best held-out loss
−0.210 lower than baseline · 95% [−0.230, −0.191]
Quality, perplexity
18.9% better
Per seed
−0.239 −0.184 −0.213 −0.203
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

Encyclopedia (near-control)

GO
Language model with 8× the capacity, 41.7M parameters
03,0006,0004.29baselinecurated
BaselineCuratedBaseline's best
Headline
22% less compute
Reached the baseline's best
step 4,690 of 5,999
Best held-out loss
−0.023 lower than baseline · 95% [−0.040, −0.011]
Quality, perplexity
2.3% better
Per seed
−0.047 −0.007 −0.018 −0.019
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

Encyclopedia, tiny sample

DON'T
Language model with 8× the capacity, 41.7M parameters
03,0006,0004.29baselinecurated
BaselineCuratedBaseline's best
Headline
never catches up
Best held-out loss
+0.063 higher than baseline · 95% [+0.038, +0.077]
Quality, perplexity
6.4% worse if curated
Per seed
+0.026 +0.078 +0.069 +0.077
Runs
4 seeds × 2 arms · 6,000 steps · 31 evaluations · 31,137 documents kept

41.7M parameters, four seeds per arm, six thousand steps. The medical gain grew, the near-control gain came out identical to the small model, and the DON'T stayed a DON'T.

2026-09-24
In progress

Image and text together

A 256M-parameter vision-language model, continued on curated image-text pairs. The first completed target of a 48-run battery, and the same target run end to end through the public service.

Charts and plots

GO
Vision-language model, 256M parameters, continued training
04,1008,3000.86baselinecurated
BaselineCuratedBaseline's best
Headline
67% less compute
Reached the baseline's best
step 2,696 of 8,249
Best held-out loss
−0.057 lower than baseline
Quality, perplexity
5.5% better
Per seed
−0.057
Runs
1 seed × 2 arms · 8,261 steps · 35 evaluations · 66,087 documents kept

Charts and plots, through the public API

GO
Same model, verdict and selection list from the service, trained on a different GPU stack
04,1008,3000.87baselinecurated
BaselineCuratedBaseline's best
Headline
65% less compute
Reached the baseline's best
step 2,922 of 8,260
Best held-out loss
−0.057 lower than baseline
Quality, perplexity
5.5% better
Per seed
−0.057
Runs
1 seed × 2 arms · 8,261 steps · 35 evaluations · 66,087 documents kept

Single seed so far; these are signs, not intervals. Through the service, the verdict was issued and sealed before training, the selection list came from the API, and the run happened on a different GPU stack and framework; its gap matched our own run of the same target to three decimals. The remaining 42 runs of the battery follow.

Scope

Every verdict was issued and timestamped before its training run began. Models up to 41.7M parameters on real text, and a 256M-parameter vision-language model in progress; larger models are the next sealed tests. Every run, including the ones we lost, sits in an audit ledger available under NDA.