Experiment results
Four questions: benchmark performance, quality curation, scaling, and better collection traces.
Saved-results snapshot · 2026-10-06 23:24 UTC
Evaluation agent: Haiku 4.5. CWM: Opus 5. Hybrid includes real checks after clean predictions and at submit; pure CWM gives no real code feedback during the episode. Lite was used for development; pool scores are not benchmark medals.
Full MLE-bench resultsQuota 58/69 valid vs Ceiling 52/69; medals 2 vs 2. One incomplete-coverage repeat.
One repeat targeting all 75 competitions. Quota hybrid uses 952 rubrics and Category-split retrieval. Ceiling executes code; Floor has no code feedback. Outcomes use the official MLE-bench grader.
| Setting | Valid | Medals | Above median |
|---|---|---|---|
| Ceiling | 52 / 6975.4% | 2 / 692.9% | 7 / 6910.1% |
| Floor | 34 / 6949.3% | 1 / 691.4% | 3 / 694.3% |
| Quota hybrid | 58 / 6984.1% | 2 / 692.9% | 4 / 695.8% |
Outside Lite: Ceiling 34/47; Floor 25/47; Quota hybrid 40/47 valid on 47 common tasks.
Coverage: Ceiling 69/75; Floor 72/75; Quota hybrid 71/75. Three grader incompatibilities and three Ceiling pod losses are excluded from the historical common cohort. Filtering favors Ceiling: Floor and Quota succeeded on two excluded tasks. This is not a complete or unbiased 75-task estimate.
Takeaway: promising validity beyond Lite, but no demonstrated medal or above-median advantage. One repeat, incomplete coverage.
Curation experiments to improve medalsWriter, evidence and retrieval changes have not established a deployable quality library or a medal gain.
Starting gap: 29 performance-quality rubrics in Quota, but 0 retrieved exposures across 924 saved control verdicts.
| Change tested | Result |
|---|---|
| Repair and structure the rubric writer | Original → repaired → structured: 1 → 0 → 4 accepted individual lessons; minimum deployable library: 12. |
| Use full traces instead of final programs | Latest teacher-specific acceptance counts are below; neither evidence format produced a deployable library. |
| Single / multiple examples / explicit counterexamples | Source-audit passes: 4 / 5 / 2; positive held-task contrasts: 0 / 0 / 0. Poor coverage prevents ranking these variants; library gates were not evaluated. |
| Reserve two retrieval slots for quality rubrics | Ordering change -14.29 pp, 95% CI [-28.57 pp, +0.00 pp]. No established positive gain on the pool screen. |
| Execute lesson-guided program revisions | Guided − unchanged: +0.72 pp balanced accuracy; 95% CI [-0.56 pp, +1.98 pp]. 5 TRAIN tasks, 15 matched repetitions—not a medal evaluation. |
Accepted lessons, source-audit passes and deployable libraries are different gates. The revision result conditions on valid matched executions; missing/failed attempts are not zero quality. None of these screens constitutes a downstream MLE-bench medal test of an accepted new library.
Takeaway: we have not yet turned quality-focused evidence into validated, transferable rubric feedback. No demonstrated medal improvement—not proof that quality curation cannot work.
Scaling library size and episode timeNo reliable gain from more rubrics or doubling the allowed episode budget in these runs.
More rubrics
Quota curation + Category-split retrieval, still capped at 12 rubrics per request. Hybrid uses real checks; pure supplies only CWM code feedback. These are separate experiments.
Hybrid · library size (not time)
Pure · library size (not time)
Hybrid: three repeats per subset, four for the original library. Pure: four repeats per size; one independently curated library per fresh size, not repeated curation. Actual library sizes are shown. No fresh-curation hybrid scaling run was completed.
More episode time
Allowed budget: 2 h / 250 steps → 4 h / 500 steps. Four baseline repeats, two longer-budget repeats.
| Setting | Mean valid / 22 | Medal rate | Above median |
|---|---|---|---|
| Ceiling | 16.50 → 16.50 | 8.0% → 6.8% | 18.2% → 15.9% |
| Quota hybrid | 18.75 → 18.00 | 4.5% → 4.5% | 9.1% → 11.4% |
Rates use all 22 tasks per repeat, including the known grader failure, not only valid submissions. Command limits stayed at 600 s for Ceiling and 180 s for hybrid checks: this did not test longer uninterrupted training.
Takeaway: these runs show no monotonic library-size trend or reliable benefit from doubling the episode budget. They do not establish that scaling cannot help.
Does a stronger trace-collection model give better rubrics? (Haiku vs. Sonnet)Sonnet's programs scored higher, but each curation run accepted at most 1 rubric; a quality library needs 12.
Question: rubrics are mined from traces of an agent solving collection-pool tasks. If a stronger model produces those traces, do we get better modeling-quality rubrics? This changes only who collects the traces; the agent evaluated on MLE-bench stays Haiku 4.5.
Setup: Haiku 4.5 and Sonnet 5.5 each made 128 attempts on the same 16 small OpenML tabular tasks. Two measurements: (1) how good the collected programs are, by held-out balanced accuracy; (2) how many quality lessons (candidate rubrics) pass curation when mined from those traces, once from the final programs only and once from the full step-by-step traces.
| Collected by | Attempts scored | Mean balanced accuracy | Lessons accepted, from final programs | Lessons accepted, from full traces |
|---|---|---|---|---|
| Haiku 4.5 | 127/128 | 85.57% | 0 | 1 |
| Sonnet 5.5 | 127/128 | 89.47% | 1 | 1 |
Accuracy averages the scored attempts. Unscored attempts: Haiku 1, Sonnet 1. Counting every attempt, Sonnet's lead is +3.09 pp to +4.65 pp depending on what those attempts would have scored (0 to 1). That is the range for the missing scores, not a confidence interval. Same time, step, command and hardware limits for both; API spend was not matched. A lesson is accepted only if it is supported by its source traces and correctly separates better from worse programs on other tasks; a quality library needs at least 12.
Takeaway: Sonnet wrote better programs, but that did not give us more usable rubrics: each curation run accepted 0 to 1 lessons, far short of the 12 needed to build a quality library and test it on an agent. The limiting step is getting lessons through curation, not the quality of the collected traces. Nothing here was run on MLE-bench.