Experiment results

Four questions: benchmark performance, quality curation, scaling, and better collection traces.

Saved-results snapshot · 2026-10-06 23:24 UTC

Evaluation agent: Haiku 4.5. CWM: Opus 5. Hybrid includes real checks after clean predictions and at submit; pure CWM gives no real code feedback during the episode. Lite was used for development; pool scores are not benchmark medals.

Full MLE-bench resultsQuota 58/69 valid vs Ceiling 52/69; medals 2 vs 2. One incomplete-coverage repeat.

One repeat targeting all 75 competitions. Quota hybrid uses 952 rubrics and Category-split retrieval. Ceiling executes code; Floor has no code feedback. Outcomes use the official MLE-bench grader.

69 commonly graded competitions · rates use every task in this cohort
SettingValidMedalsAbove median
Ceiling52 / 6975.4%2 / 692.9%7 / 6910.1%
Floor34 / 6949.3%1 / 691.4%3 / 694.3%
Quota hybrid58 / 6984.1%2 / 692.9%4 / 695.8%

Outside Lite: Ceiling 34/47; Floor 25/47; Quota hybrid 40/47 valid on 47 common tasks.

Coverage: Ceiling 69/75; Floor 72/75; Quota hybrid 71/75. Three grader incompatibilities and three Ceiling pod losses are excluded from the historical common cohort. Filtering favors Ceiling: Floor and Quota succeeded on two excluded tasks. This is not a complete or unbiased 75-task estimate.

Takeaway: promising validity beyond Lite, but no demonstrated medal or above-median advantage. One repeat, incomplete coverage.

Curation experiments to improve medalsWriter, evidence and retrieval changes have not established a deployable quality library or a medal gain.

Starting gap: 29 performance-quality rubrics in Quota, but 0 retrieved exposures across 924 saved control verdicts.

Attempts to improve modeling quality · collection-pool studies
Change testedResult
Repair and structure the rubric writerOriginal → repaired → structured: 1 → 0 → 4 accepted individual lessons; minimum deployable library: 12.
Use full traces instead of final programsLatest teacher-specific acceptance counts are below; neither evidence format produced a deployable library.
Single / multiple examples / explicit counterexamplesSource-audit passes: 4 / 5 / 2; positive held-task contrasts: 0 / 0 / 0. Poor coverage prevents ranking these variants; library gates were not evaluated.
Reserve two retrieval slots for quality rubricsOrdering change -14.29 pp, 95% CI [-28.57 pp, +0.00 pp]. No established positive gain on the pool screen.
Execute lesson-guided program revisionsGuided − unchanged: +0.72 pp balanced accuracy; 95% CI [-0.56 pp, +1.98 pp]. 5 TRAIN tasks, 15 matched repetitions—not a medal evaluation.

Accepted lessons, source-audit passes and deployable libraries are different gates. The revision result conditions on valid matched executions; missing/failed attempts are not zero quality. None of these screens constitutes a downstream MLE-bench medal test of an accepted new library.

Takeaway: we have not yet turned quality-focused evidence into validated, transferable rubric feedback. No demonstrated medal improvement—not proof that quality curation cannot work.

Scaling library size and episode timeNo reliable gain from more rubrics or doubling the allowed episode budget in these runs.

More rubrics

Quota curation + Category-split retrieval, still capped at 12 rubrics per request. Hybrid uses real checks; pure supplies only CWM code feedback. These are separate experiments.

Hybrid · library size (not time)

Hybrid · library size (not time): valid submissions by actual rubric countVertical range is fixed from 0 to 22 valid submissions. Each hollow circle is one complete repeat; the solid bar is the mean. Sizes are equally spaced categories, not a linear numeric axis. Incomplete groups have no plotted points or mean. Exact repeat values are available in each point's tooltip and the aggregate download.Valid submissions / 22 tasks051015202210Nested subset · nominal 10; actual N=10; repeat 1: 19/22 validNested subset · nominal 10; actual N=10; repeat 2: 16/22 validNested subset · nominal 10; actual N=10; repeat 3: 19/22 validNested subset · nominal 10: mean 18.00/22 over 3 repeats3 repeats18.00100Nested subset · nominal 100; actual N=100; repeat 1: 16/22 validNested subset · nominal 100; actual N=100; repeat 2: 14/22 validNested subset · nominal 100; actual N=100; repeat 3: 16/22 validNested subset · nominal 100: mean 15.33/22 over 3 repeats3 repeats15.33500Nested subset · nominal 500; actual N=500; repeat 1: 16/22 validNested subset · nominal 500; actual N=500; repeat 2: 19/22 validNested subset · nominal 500; actual N=500; repeat 3: 18/22 validNested subset · nominal 500: mean 17.67/22 over 3 repeats3 repeats17.67952Original full library; actual N=952; repeat 1: 20/22 validOriginal full library; actual N=952; repeat 2: 16/22 validOriginal full library; actual N=952; repeat 3: 20/22 validOriginal full library; actual N=952; repeat 4: 19/22 validOriginal full library: mean 18.75/22 over 4 repeats4 repeats18.75Actual rubric count (categorical spacing)
Points: repeats; bars and labels: means out of 22 Lite tasks.

Pure · library size (not time)

Pure · library size (not time): valid submissions by actual rubric countVertical range is fixed from 0 to 22 valid submissions. Each hollow circle is one complete repeat; the solid bar is the mean. Sizes are equally spaced categories, not a linear numeric axis. Incomplete groups have no plotted points or mean. Exact repeat values are available in each point's tooltip and the aggregate download.Valid submissions / 22 tasks05101520228Fresh curation · nominal 10; actual N=8; repeat 1: 6/22 validFresh curation · nominal 10; actual N=8; repeat 2: 8/22 validFresh curation · nominal 10; actual N=8; repeat 3: 7/22 validFresh curation · nominal 10; actual N=8; repeat 4: 10/22 validFresh curation · nominal 10: mean 7.75/22 over 4 repeats4 repeats7.7599Fresh curation · nominal 100; actual N=99; repeat 1: 12/22 validFresh curation · nominal 100; actual N=99; repeat 2: 12/22 validFresh curation · nominal 100; actual N=99; repeat 3: 12/22 validFresh curation · nominal 100; actual N=99; repeat 4: 13/22 validFresh curation · nominal 100: mean 12.25/22 over 4 repeats4 repeats12.25500Fresh curation · nominal 500; actual N=500; repeat 1: 6/22 validFresh curation · nominal 500; actual N=500; repeat 2: 8/22 validFresh curation · nominal 500; actual N=500; repeat 3: 11/22 validFresh curation · nominal 500; actual N=500; repeat 4: 8/22 validFresh curation · nominal 500: mean 8.25/22 over 4 repeats4 repeats8.25952*Reused original anchor · nominal 1,000; actual N=952; repeat 1: 11/22 validReused original anchor · nominal 1,000; actual N=952; repeat 2: 11/22 validReused original anchor · nominal 1,000; actual N=952; repeat 3: 13/22 validReused original anchor · nominal 1,000; actual N=952; repeat 4: 12/22 validReused original anchor · nominal 1,000: mean 11.75/22 over 4 repeats4 repeats11.751,802Fresh curation · nominal 2,000; actual N=1802; repeat 1: 11/22 validFresh curation · nominal 2,000; actual N=1802; repeat 2: 11/22 validFresh curation · nominal 2,000; actual N=1802; repeat 3: 12/22 validFresh curation · nominal 2,000; actual N=1802; repeat 4: 14/22 validFresh curation · nominal 2,000: mean 12.00/22 over 4 repeats4 repeats12.00Actual rubric count (categorical spacing)
Points: repeats; bars and labels: means out of 22 Lite tasks. * 952: reused original library, not a freshly curated size.

Hybrid: three repeats per subset, four for the original library. Pure: four repeats per size; one independently curated library per fresh size, not repeated curation. Actual library sizes are shown. No fresh-curation hybrid scaling run was completed.

More episode time

Allowed budget: 2 h / 250 steps → 4 h / 500 steps. Four baseline repeats, two longer-budget repeats.

2-hour → 4-hour budget results · MLE-bench Lite
SettingMean valid / 22Medal rateAbove median
Ceiling16.50 → 16.508.0% → 6.8%18.2% → 15.9%
Quota hybrid18.75 → 18.004.5% → 4.5%9.1% → 11.4%

Rates use all 22 tasks per repeat, including the known grader failure, not only valid submissions. Command limits stayed at 600 s for Ceiling and 180 s for hybrid checks: this did not test longer uninterrupted training.

Takeaway: these runs show no monotonic library-size trend or reliable benefit from doubling the episode budget. They do not establish that scaling cannot help.

Does a stronger trace-collection model give better rubrics? (Haiku vs. Sonnet)Sonnet's programs scored higher, but each curation run accepted at most 1 rubric; a quality library needs 12.

Question: rubrics are mined from traces of an agent solving collection-pool tasks. If a stronger model produces those traces, do we get better modeling-quality rubrics? This changes only who collects the traces; the agent evaluated on MLE-bench stays Haiku 4.5.

Setup: Haiku 4.5 and Sonnet 5.5 each made 128 attempts on the same 16 small OpenML tabular tasks. Two measurements: (1) how good the collected programs are, by held-out balanced accuracy; (2) how many quality lessons (candidate rubrics) pass curation when mined from those traces, once from the final programs only and once from the full step-by-step traces.

Trace-collection model · same 16 pool tasks
Collected byAttempts scoredMean balanced accuracyLessons accepted, from final programsLessons accepted, from full traces
Haiku 4.5127/12885.57%01
Sonnet 5.5127/12889.47%11

Accuracy averages the scored attempts. Unscored attempts: Haiku 1, Sonnet 1. Counting every attempt, Sonnet's lead is +3.09 pp to +4.65 pp depending on what those attempts would have scored (0 to 1). That is the range for the missing scores, not a confidence interval. Same time, step, command and hardware limits for both; API spend was not matched. A lesson is accepted only if it is supported by its source traces and correctly separates better from worse programs on other tasks; a quality library needs at least 12.

Takeaway: Sonnet wrote better programs, but that did not give us more usable rubrics: each curation run accepted 0 to 1 lessons, far short of the 12 needed to build a quality library and test it on an agent. The limiting step is getting lessons through curation, not the quality of the collected traces. Nothing here was run on MLE-bench.