One clean run · final split ratio_40_60_final · generated 2026-08-18
Experiment setting
Competition pool: MLE-bench ships 75 competitions; 14 are excluded — 9 with broken/unpreparable dataset manifests, 2 with Kaggle checksum drift (h-and-m, icecube: official grading ground truth can no longer be reconstructed), and 3 runtime outliers whose collection would take >2× the next-slowest competition (facebook-recruiting, billion-word, vinbigdata) — leaving 61 usable.
Split: 24 mining / 37 held-out evaluation competitions (~40% : 60% of the 61 usable). Stages are nested prefixes of this one split — collect once, judge once, slice per stage.
Collection turns (per program): 1 generation turn (haiku) → if it crashes, up to 3 repair turns (each crash harvested); if it runs clean, official grading then up to 8 climb turns (sonnet-class improver, each version executed + graded).
Collection budget (per competition): up to 40 programs; stop at 15 curated error rubrics + 15 perf candidates, or after 8 consecutive programs yielding nothing new (stagnation); 3600s execution cap, one program per GPU.
Curation: inline v13 error curation (exception-class committed) with universal-template lint; perf pairs require a measured score improvement above the competition's replicate-noise floor.
Judge: claude-opus-5 (fast mode), blind — sees one rubric + task + program code only, no execution info.
Perf truth: adjudicated annotator credits against measured weak-vs-strong score gaps (mechanism-level match counts).
Controls: generic and shuffled rubrics score 0.000 precision at every stage.
Final library: 381 error rubrics, 327 perf rubrics.
Error R = per-program catch rate (≥1 correct-class fire) over catchable failures. Full-scale CIs (10k bootstrap over trajectories):
error P [0.31–0.46] · perf P [0.19–0.33] · perf R [0.34–0.49].