Rubric libraries vs the errors the agent actually makes (MLE-bench)
The MLE rubric libraries, how each was built, their error mix against the collection pool and MLE-bench Lite, what the judge was given, the results, and the harness fixes · built 2026-09-29 00:43 · index
Status, 2026-09-28.
Categories made finer (done). One category, runtime/environment failure, held 90% of MLE-bench's failures. It is now split into failure types for failures and rubrics alike (section 2).
Filtered oracle pool and filtered oracle built on the finer categories (done). Filtered oracle pool (464 rubrics): 1 from the collection pool, but 42 from MLE-bench, because the pool itself is 42 from MLE-bench (time limits are 18% of the pool's failures, 56% of MLE-bench's). Oracle (177 rubrics): 2 from MLE-bench, time limits 58% vs 56%. It is small because the whole-rollout library has only 102 time-limit rubrics. Earlier versions of both, matched on the 21 coarse categories only, were run and scored like the other libraries; they were not real matches and are no longer shown.
First five libraries run with the final harness (2026-09-27, section 3). wrong final answers 12; whole rollouts 11; oracle pool 17; oracle 16; weighted 17 / 16; against 18 with every command running and 7 with no feedback (valid submissions of 22).
Second round (2026-09-28, two runs each, section 3). Complete target (failures in the trace + causes found at grading), outcome weight, and twice the rollouts (second + third collections, 3,550 rubrics). Filtered oracle pool on the complete target 20 / 17, the only library at the level of running every command (18); weighted 14 / 16; the same library with no matching 14; filtered oracle on the complete target 14 / 17. Matching the pool beats matching MLE-bench's own mix again, and the outcome weight did not beat the first weighted library. The two matched libraries keep only 187 and 186 of the 3,550 rubrics: the library's mix is nearly the opposite of the failures' mix (section 2, 'Why only about 5%').
A harness loss found and fixed (2026-09-28). Some final runs were cut off (the log lacks the exit line the harness always writes; several runners lost theirs in the same second, pods still up), leaving an empty submission graded as the agent's failure: 2-4 competitions per run, also in the first-round runs. Such a competition is now re-run once. Re-run that way: every section-3 row except the weighted (b) runs (first-round valid counts moved by 0-2; the two reference runs had none). Not re-run: the (b) runs and the section-4 runs, which have 1-10 such competitions each, some of them possibly the agent's own out-of-memory.
Collection widened (done, not run). The next collection runs all six usable benchmarks together (section 0), which should add time-limit failures from real training jobs in MLGym and MLAgentBench.
0. The collection pool
What the collection pool is
Tasks. Every task from a set of public benchmarks that are not MLE-bench (listed below), after a filter removes any task that is an MLE-bench competition or shares its distinctive name words (all 75 MLE-bench competitions plus the dev one, so all 22 of MLE-bench Lite; rubrics_gym/holdouts/mle-bench.yaml). Nothing is sampled: every task that survives the filter and has its data uploaded is used.
Rollouts. The same Haiku agent that is evaluated attempts each task several times (a rollout is one attempt): 4 times for the small one-CSV question tasks, 8 times for the competition-style tasks, whose limits copy MLE-bench's (2 h, 250 steps, $3, 600 s per command, 3 CPU / 11 GB).
Recording. Every command the agent runs and every output it gets back is saved.
What comes out. (a) Opus writes rubrics from those recordings: that is every library below. (b) The failed commands in them are the pool's error distribution, the blue bars in section 2: what the agent gets wrong on tasks like MLE-bench's without being on MLE-bench.
So the pool is not a fixed dataset: it is whichever collection run produced the rubrics. The one behind the first-round libraries and their charts is the third collection (2026-09-23), below; the second-round libraries (2026-09-28) were written from the second and third collections together (2,475 rollouts, same design and mix) and are compared with their complete target (section 2).
benchmark (third collection)
tasks
rollouts
share of rollouts
failed commands
share of the pool's failures
InfiAgent-DABench
217
868
70%
135
19%
DA-Code
53
212
17%
169
24%
DSBench
21
168
13%
401
57%
total
291
1,248
705
DSBench is 13% of the rollouts but 57% of the failures and 95% of the time-limit failures: the one-CSV tasks almost never hit a limit. The next collection adds MLGym, MLAgentBench and RE-Bench (table below: 318 tasks, 1,464 rollouts), so the pool behind the next libraries will have a different mix.
Every benchmark the pool can draw from
Status of each benchmark, the filter's effect, and how many rollouts each collection took from it.
benchmark
what the tasks are
tasks
kept after the filter
rollouts, first collection
rollouts, second collection
rollouts, third collection
ready for the next collection
InfiAgent-DABench
Data-analysis questions over one CSV, with a checkable numeric or text answer.
257
217 (40 removed)
4,337
859
868
all kept tasks
DA-Code
Data-science tasks (wrangling, analysis, ML, plots) over provided files; only the 62 whose data ships in the repository are usable.
500
56 (6 removed)
1,060
210
212
all kept tasks
DSBench
Kaggle-style modelling tasks: train a model on the provided data and write a submission file, scored by the competition metric.
75
21 (54 removed)
84
158
168
21 tasks
MLGym
ML research tasks: train and improve a model against a baseline (game theory, RL, image, text, tabular).
19
13 (1 removed)
–
–
–
13 tasks
MLAgentBench
ML experimentation tasks: improve a given training script's metric.
15
12 (0 removed)
–
–
–
12 tasks
RE-Bench
ML research-engineering tasks; only 2 of 7 fit one small GPU.
7
2 (0 removed)
–
–
–
2 tasks
DABstep
Multi-step data-analysis questions over payment data.
450
–
–
–
–
not usable: answers hidden
Until now DSBench was the only competition-style source: MLGym, MLAgentBench and RE-Bench had loaders but their data was never uploaded, so no rollout from them exists and all five libraries below come from InfiAgent-DABench, DA-Code and 21 DSBench tasks. From 2026-09-27 a collection runs every usable benchmark together by default (mining/run_da_pool.py, 4 rollouts per InfiAgent-DABench / DA-Code task, 8 per competition or research task, CPU-only like the MLE-bench runs) and refuses to start if a benchmark is missing. MLAgentBench tasks whose data needs a manual step or a Kaggle competition the account has not joined stay out until prepared.
collection
date
rollouts
libraries built from it
First collection
2026-09-09 / 09-21
5,481
rubrics from wrong final answers (its InfiAgent-DABench + DA-Code part)
Second collection
2026-09-22
1,227
with the third: the larger weighted library and its oracle pool / oracle (2026-09-28)
Third collection
2026-09-23
1,248
rubrics from whole rollouts (4,294) and the three libraries derived from them; with the second, the 2026-09-28 libraries
1. The rubric libraries and how each was implemented
Every library starts from the same place: the rollouts of the collection pool (section 0; the wrong-answer library from the first collection, all others from the third). What differs is what Opus reads when it writes rubrics, and whether anything is done to the library afterwards.
The two baselines
Wrong final answers (the original): only rollouts that ended with a wrong answer, and only their final scripts and answer, so it knows nothing about the crashes on the way. Whole rollouts, no matching: every state of every rollout, everything Opus writes kept (near-duplicates removed). The oracle pool and the oracle are filtered from this library.
Distribution matching: oracle pool and oracle
A filter, not new mining. Source: the 4,294 rubrics of the whole-rollout library, each sorted into one failure category (section 2). Steps: (1) take the target, the failures of the collection pool (oracle pool) or of MLE-bench Lite (oracle), and compute each category's share; (2) give each category a budget of round(total x share) rubrics; (3) walk the whole-rollout library in the order the rubrics were written and keep a rubric while its category has budget left. The total is the largest at which every category's budget can be filled, so no category runs short: for the oracle, time limits are 56.2% of MLE-bench's failures and the library has 102 time-limit rubrics, so the total is 177 and time-limit rubrics are 57.6% of it. Categories the library has no rubric for (disk full; out of memory for the pool) are left out of the target. The oracle looks at the test set, so it is a ceiling check, not a method.
Weighted
New mining of the same rollouts, nothing filtered. Every candidate rubric must cite the step numbers it comes from; its weight is how many of those steps were errors. Candidates saying the same thing are merged and their weights added (4,638 candidates, 2,276 rubrics; 1,820 weigh 0). The whole library is kept; the weight only changes which 12 rubrics the judge is given (the two rules in its card). Checked: for every rubric whose members are all listed (2,267), its weight equals the sum of its members' weights recomputed from the trajectories.
Check: does each match hit its target?
library
its target
difference to its target
difference to MLE-bench
Filtered oracle pool, first round
the third collection's trace failures
0.5
42
Filtered oracle, first round
MLE-bench Lite's trace failures
2.2
2
Filtered oracle pool (complete target)
2nd+3rd collections: trace failures + causes at grading
1.6
50
Filtered oracle (complete target)
MLE-bench Lite: trace failures + causes at grading
2.1
2
0 = identical, 100 = no overlap. All four hit their target (the first two recomputed independently from the raw failure records). The oracle pools are still far from MLE-bench because the pool is.
Second round (2026-09-28): complete target, outcome weight, twice the rollouts
Complete target. The targets above count only commands that failed. The complete target adds, for every rollout that ended wrong, the cause found at grading (wrong format, output never written, no validation...), in the same categories (section 2). It is 31% causes at grading for the pool and 13% for MLE-bench.
Outcome weight. A weighted rubric also gets 1 for each rollout that ended wrong whose final scripts or answer it cites; wrong-result rubrics stop weighing 0 (outcome weight is 1,974 of the library's 4,207).
Twice the rollouts. The second collection (1,227 rollouts, same design as the third) mined with the same miner and pooled with the third: 3,550 rubrics (726 crash, 2,572 wrong-result, 252 crash-caused wrong-result).
Filtered oracle pool and filtered oracle rebuilt from that library on the complete targets: 187 and 186 rubrics. They did not grow: the filter needs every category's budget to fill, and the whole-rollout miner writes few missing-package (29) and time-limit (93) rubrics, which are 15% of the pool's and 50% of MLE-bench's complete target. The oracle now has 4 wrong-format and 19 no-validation rubrics, the causes of most of MLE-bench's wrong results; the earlier oracle had none of either, and its detecting-insults submission came out with one of three columns.
Each library end to end
library
rubrics come from
what Opus reads
what Opus writes
clean-up
after mining
how the judge gets 12 at run time
rubrics
valid of 22
Wrong final answers (baseline)
First collection (2026-09-09): 5,397 rollouts of InfiAgent-DABench + DA-Code (about 20 per task); only rollouts whose final answer was graded wrong are read
The task, the final scripts, the submitted answer and the grader's verdict. Not the commands along the way or their output.
One rubric per wrong rollout: what a reviewer could have checked in the scripts to catch the wrong answer.
Near-duplicates dropped (embedding cosine 0.90).
Nothing.
12 most similar to the task and the code change
1,076
12
Whole rollouts, no matching (baseline)
Third collection (2026-09-23): 1,248 rollouts = DSBench 21 tasks x 8, InfiAgent-DABench 217 x 4, DA-Code 53 x 4
The whole trace: every command and its output, each crash marked (and whether the agent recovered), the final scripts, the answer and the verdict.
Up to 6 rubrics per rollout, each tagged crash / wrong-result / crash-caused wrong-result: 5,232 candidates.
Third collection (2026-09-23): 1,248 rollouts = DSBench 21 tasks x 8, InfiAgent-DABench 217 x 4, DA-Code 53 x 4 (the same rollouts, mined a second time)
The whole trace (up to 400k characters); every rubric must cite the step numbers it comes from.
Up to 6 rubrics per rollout: 4,638 candidates. Weight = how many cited steps were errors (82% weigh 0).
Rubrics saying the same thing merged, weights added (embedding cosine 0.82): 2,276.
Nothing dropped.
(a) 12 split between the three kinds in proportion to total weight, most similar first plus a small weight bonus; (b) only weight > 0, split by weight squared
2,276
17 (a) / 13 (b)
Filtered oracle pool, first round (matched to the pool's failed commands)
Library 'whole rollouts' (its 4,294 rubrics); target = the failures of the same 1,248 rollouts
Nothing new: a filter.
Nothing new.
–
Budget per failure category = round(total x the pool's share); rubrics kept in writing order until the budget is full; total = the largest at which every budget fills.
12 most similar to the task and the code change
464
17
Filtered oracle, first round (matched to MLE-bench's failed commands)
Library 'whole rollouts' (its 4,294 rubrics); target = MLE-bench Lite's failures (the test set)
Nothing new: a filter.
Nothing new.
–
Same budget filter with MLE-bench's shares; the 102 time-limit rubrics cap the total.
12 most similar to the task and the code change
177
16
Weighted, second + third collections
Second + third collections (2026-09-22, 09-23): 2,475 rollouts, the same design twice
The whole trace; every rubric must cite its steps (third collection's candidates reused, second mined new).
Up to 6 rubrics per rollout: 9,094 candidates. Weight = cited steps that were errors + 1 if it cites the final scripts or answer of a rollout that ended wrong.
Rubrics saying the same thing merged, weights added (embedding cosine 0.82): 3,550.
Nothing dropped.
(a) 12 split between the three kinds in proportion to total weight; also run once with the plain 12 most similar
3,550
14 / 16
Filtered oracle pool (complete target)
Library 'weighted, 2nd+3rd' (3,550); target = the complete failures of the same 2,475 rollouts
Nothing new: a filter.
Nothing new.
–
Same budget filter, complete target (trace failures + causes found at grading); 29 missing-package rubrics cap the total.
12 most similar to the task and the code change
187
20 / 17
Filtered oracle (complete target)
Library 'weighted, 2nd+3rd' (3,550); target = MLE-bench Lite's complete failures (the test set)
Nothing new: a filter.
Nothing new.
–
Same budget filter with MLE-bench's complete shares; 93 time-limit rubrics cap the total.
12 most similar to the task and the code change
186
14 / 17
Common to every library: the rollouts come from the same Haiku agent that is evaluated, on benchmark tasks filtered against MLE-bench (section 0); Opus writes every rubric in the same shape (title, when it applies, pattern, how to detect it from the task and code, what is not a violation, consequence); Haiku sorts every rubric into one failure category (section 2). At run time the judge (Opus) sees the 12 rubrics, the task, the code change, a read-only view of the sandbox and data and the last real run; the agent sees only a predicted reward and each applicable rubric with its own 0/1 (section 4). Valid of 22: one run each with the final harness (weighted: first run of each picking rule; second-round libraries: run 1 / run 2).
Each library in detail
The oracle pool and the oracle load from local files (not on Hugging Face).
library
rubrics
kinds
made by
Wrong final answers (baseline) mle-rubrics-onpolicy-all
1,076
1,076 from wrong answers
written from rollouts
Whole rollouts, no matching (baseline) mle-rubrics-clean
filtered from the 2nd+3rd library, complete target (uses the test set)
Wrong final answers (baseline) 1,076 rubricsmle-rubrics-onpolicy-all
How it was built
First collection, InfiAgent-DABench + DA-Code only (5,397 rollouts of one-CSV questions). For every rollout whose final answer was graded wrong, Opus reads the task, the final scripts, the answer and the grader's verdict, and writes one rubric: what a reviewer could have checked in the scripts to catch it. It never sees the commands the agent ran or their errors, so it cannot write about crashes.
What is kept
Only wrong final answers are read. A rubric is dropped when its title and applies-when line are at least 0.90 cosine-similar (text embeddings) to one already kept. No matching to any error distribution.
Whole rollouts, no matching (baseline) 4,294 rubricsmle-rubrics-clean
How it was built
Third collection: 1,248 rollouts, 168 DSBench (21 tasks x 8) and 1,080 InfiAgent-DABench + DA-Code (270 tasks x 4). One Opus call per rollout with the whole trace: every command and its output, each crash marked (and whether the agent recovered), the final scripts, the submitted answer and the grader's verdict. Opus writes up to 6 rubrics per rollout, each tagged as a crash rubric (a code pattern that made a run die), a wrong-result rubric (the code ran but the answer was wrong) or a crash-caused wrong-result rubric (a crash forced a fallback that made the answer wrong). 5,232 candidate rubrics.
What is kept
Everything Opus writes is kept except word-overlap duplicates (TF-IDF cosine 0.50), which removed 18%. No matching. A rollout supports several wrong-result rubrics but only one or two crash rubrics, so the library ends up 68% wrong-result rubrics while most of the agent's errors are crashes. The oracle pool and the oracle are filtered from this library.
Same 1,248 rollouts, mined again. Opus reads the full trace (up to 400k characters) and every rubric must cite the step numbers it comes from. A rubric's weight is how many of its cited steps were errors (non-zero exit, Traceback, Killed, timed out, MemoryError). 82% of candidates cite no error step and weigh 0, including 96% of wrong-result rubrics. Rubrics that say the same thing are merged instead of dropped (0.82 cosine, then again between merged groups) and their weights added: 4,638 candidates become 2,276 rubrics.
What is kept
Nothing is dropped; how often each problem happened is kept as the weight and used when choosing which 12 rubrics the judge gets. Two ways of choosing were run: (a) the 12 are split between the three kinds in proportion to each kind's total weight (about 10 crash, 1 crash-caused, 1 wrong-result), most similar to the task first with a small bonus for weight; (b) only rubrics with weight above 0 can be chosen, split in proportion to weight squared, so all 12 are crash rubrics. The purple bars count each rubric by its weight.
Filtered oracle pool, first round (matched to the pool's failed commands) 464 rubricsmle-rubrics-clean-m2split-fit
How it was built
A filter of the whole-rollout library; nothing new is written. Target: the failures of the collection pool the library was written from (all 1,248 rollouts). Per category, round(total x its share of the pool's failures) rubrics, taken in the order they were written. The total is the largest at which every category's budget can be filled, so no category runs short.
What is kept
Run once with the final harness (section 3); local file mle-rubrics-clean-m2split-fit, not on Hugging Face.
Code
mining/curate_split.py
Filtered oracle, first round (matched to MLE-bench's failed commands) 177 rubricsmle-rubrics-clean-oracledist-split-fit
How it was built
The same filter with MLE-bench Lite's own failures as the target. Time limits are 56% of them and the whole-rollout library has only 102 time-limit rubrics, so the largest total at which every budget fills is small. Categories no rubric covers (disk full) are left out. It looks at the test set, so it is a ceiling check, not a method.
What is kept
Run once with the final harness (section 3); local file mle-rubrics-clean-oracledist-split-fit, not on Hugging Face.
Code
mining/curate_split.py
Weighted, second + third collections 3,550 rubricsmle-rubrics-clean-v3
How it was built
The weighted library's miner run on the second collection too (2026-09-22, 1,227 rollouts of the same design: 158 DSBench, 1,069 InfiAgent-DABench + DA-Code), pooled with the third (the third's cached candidates reused): 2,475 rollouts, 9,094 candidates, merged into 3,550 rubrics. The weight now has two parts: error weight, as before (cited steps that were errors), plus outcome weight, 1 when the rubric cites the final scripts or answer of a rollout that ended wrong. Before, a wrong answer never produced an error step, so 96% of wrong-result rubrics weighed 0; now 1,974 of the 4,207 weight comes from wrong outcomes.
What is kept
Nothing is dropped. Run with picking rule (a) (12 split between the kinds in proportion to total weight) twice, and once with the plain 12 most similar and no weight (the no-matching baseline for this source). Local file, not on Hugging Face.
Filtered oracle pool (complete target) 187 rubricsmle-rubrics-clean-v3-oraclepool
How it was built
The budget filter on the 2nd+3rd library, with a complete target: the failures in the trace of all 2,475 rollouts plus, for each rollout that ended wrong, the cause found at grading (section 2). Missing-package failures are 15% of that target and the library has 29 missing-package rubrics, so the total is small.
What is kept
Run twice with the final harness. Local file, not on Hugging Face.
The same filter with MLE-bench Lite's complete target: its trace failures plus the cause of each of the 19 competitions the interpreter run got wrong (no valid submission, or below the median). Time limits are half the target and the library has 93 time-limit rubrics. Looks at the test set: a ceiling check, not a method.
What is kept
Run twice with the final harness. Local file, not on Hugging Face.
Code
mining/curate_complete.py
Every rubric has the same shape: a title, when it applies, the pattern, how a reviewer detects it from the task and the code alone (the judge never sees program output), what separates a real violation from something that only looks like one, and the consequence.
2. Error categories: the agent's errors vs each library
Errors. A code-running command that failed: non-zero exit, or an error word such as Traceback, Killed or timed out (commands that exited 0 and only printed such a word are not counted). Two sources: the collection pool the libraries were written from, all 1,248 rollouts of the third collection (705 errors), and MLE-bench Lite, the held-out test (128 errors from the run where every command executes, 22 competitions; only measured, and used to choose nothing except the oracle).
Categories. MLE-bench's 21 error categories (built once by Sonnet from MLE-bench's real-execution outputs and the wrong-answer rubrics, then frozen; Haiku sorts every error and rubric into one). One of them, Runtime/environment failure prevents execution, held 90% of MLE-bench's errors, too coarse to compare anything, so it is split here into failure types (red labels): an error gets its type from the same fixed regular expressions the crash rubrics are built on (exit code and output, no model), and a rubric in that category gets one from Haiku reading it against the same type definitions (analysis/split_runtime.py, 1997 rubrics). Inside it MLE-bench's errors are 56% time limits, 16% exceptions in the code's own logic, 9% missing packages and 6% library API changes (% of all errors).
Difference between two distributions (total variation distance): half the sum of the absolute percentage differences over the categories, the share that would have to move to turn one into the other. 0 = identical, 100 = no overlap. Shown on the split categories; the summary also gives it on the original 21, which is what the subsampled libraries were matched on.
Libraries from the second round (2nd+3rd collections) are compared with the complete targets described in the next subsection; the earlier ones with the failures in the trace of the third collection.
Wrong final answers (baseline)
collection pool errors (705)MLE-bench Lite errors (128)library (1,076)
Difference: vs collection pool errors 91, vs MLE-bench Lite errors 91. Runtime share: 0% of the library vs 82% of the pool's and 90% of MLE-bench's.
Whole rollouts, no matching (baseline)
collection pool errors (705)MLE-bench Lite errors (128)library (4,294)
Difference: vs collection pool errors 69, vs MLE-bench Lite errors 77. Runtime share: 13% of the library vs 82% of the pool's and 90% of MLE-bench's.
Weighted
collection pool errors (705)MLE-bench Lite errors (128)library (2,276)
Difference: vs collection pool errors 71, vs MLE-bench Lite errors 79. Runtime share: 12% of the library vs 82% of the pool's and 90% of MLE-bench's.
Weighted, counted by weight
collection pool errors (705)MLE-bench Lite errors (128)library, counted by weight (2,276)
Difference: vs collection pool errors 38, vs MLE-bench Lite errors 49. Runtime share: 52% of the library vs 82% of the pool's and 90% of MLE-bench's.
Filtered oracle pool, first round (matched to the pool's failed commands)
collection pool errors (705)MLE-bench Lite errors (128)library (464)
Difference: vs collection pool errors 1, vs MLE-bench Lite errors 42. Runtime share: 82% of the library vs 82% of the pool's and 90% of MLE-bench's.
Filtered oracle, first round (matched to MLE-bench's failed commands)
collection pool errors (705)MLE-bench Lite errors (128)library (177)
Difference: vs collection pool errors 42, vs MLE-bench Lite errors 2. Runtime share: 90% of the library vs 82% of the pool's and 90% of MLE-bench's.
Weighted, second + third collections
collection pool, complete target (2,018)MLE-bench Lite, complete target (147)library (3,550)
Difference: vs collection pool, complete target 56, vs MLE-bench Lite, complete target 68. Runtime share: 11% of the library vs 56% of the pool's and 78% of MLE-bench's.
Weighted, second + third collections, counted by weight
collection pool, complete target (2,018)MLE-bench Lite, complete target (147)library, counted by weight (3,550)
Difference: vs collection pool, complete target 41, vs MLE-bench Lite, complete target 57. Runtime share: 22% of the library vs 56% of the pool's and 78% of MLE-bench's.
Filtered oracle pool (complete target)
collection pool, complete target (2,018)MLE-bench Lite, complete target (147)library (187)
Difference: vs collection pool, complete target 2, vs MLE-bench Lite, complete target 50. Runtime share: 57% of the library vs 56% of the pool's and 78% of MLE-bench's.
Filtered oracle (complete target)
collection pool, complete target (2,018)MLE-bench Lite, complete target (147)library (186)
Difference: vs collection pool, complete target 49, vs MLE-bench Lite, complete target 2. Runtime share: 78% of the library vs 56% of the pool's and 78% of MLE-bench's.
The complete target: failures in the trace plus causes found at grading
The error counts above see only commands whose output showed an error. A rollout whose scripts ran cleanly but wrote the wrong columns, skipped validation or never wrote the output file adds nothing to them, although that is why it was graded wrong. For every rollout that ended wrong, Haiku reads the task, the final scripts, the submission or answer, the grader's verdict, the failures recorded along the way and (MLE-bench) the final run's log, and names the one category that explains the wrong result, in the same categories (analysis/outcome_causes.py). Ended wrong means: collection pool, the answer misses the gold value (InfiAgent-DABench, DA-Code) or there is no submission beating a constant prediction (DSBench); MLE-bench Lite, no valid submission or one below the median. Each cause counts once, next to the trace failures.
Collection pool (2nd+3rd collections, 2,475 rollouts): 1,393 trace failures + 625 causes at grading (31% of the complete target). MLE-bench Lite: 128 + 19 (13%). The pool's causes are mostly an output file never written or overwritten (48% of them) and wrong output format (14%); MLE-bench's are mostly predictions never checked on held-out data (79%) and wrong output format (16%). In the trace these categories are 0.2% + 0.1% of the pool's failures and 0.0% + 0.0% of MLE-bench's.
Trace failures only vs complete target
collection pool, trace only (1,393)collection pool, complete (2,018)MLE-bench, trace only (128)MLE-bench, complete (147)
Collection pool vs MLE-bench: trace only 44, complete 49.
Why only about 5% of the 3,550 rubrics survive the matching
The library and the agent's failures have almost opposite mixes. Runtime failures (red labels) are 56% of the collection pool's complete target and 78% of MLE-bench's, but 11% of the library: the miner mostly writes rubrics about wrong results, and a rollout gives it many of those but only one or two about the crash. Matching by filtering can only remove rubrics, so the library shrinks until its scarcest needed category is used up. Adding the second collection doubled the library but added the same mix, so the scarce categories stayed scarce (29 missing-package rubrics, 93 time-limit rubrics).
The filter keeps each category at its share of the target. A library of size T needs T × share rubrics in every category, so T can be no larger than available ÷ share for any of them: T = min over categories of (available ÷ share). For the oracle pool the smallest is Runtime: missing package: 15.4% of the target, but only 29 of the 3,550 rubrics (0.8%), so T = 29 ÷ 0.154 = 188 (187 after rounding each category's budget). Every other category is then cut to its share of 188: model output not validated against baseline has 552 rubrics and keeps 3, statistical method/definition substituted unchecked has 379 rubrics and keeps 2, parameter/threshold chosen without validation has 179 rubrics and keeps 0.
share of the collection pool's failures (the target)share of the 3,550-rubric library
How big a matched library each category could supply (rubrics available ÷ the category's share of the target); the dashed line is the smallest, 188, which becomes the library's size. Only the ten smallest are shown.
The filter keeps each category at its share of the target. A library of size T needs T × share rubrics in every category, so T can be no larger than available ÷ share for any of them: T = min over categories of (available ÷ share). For the oracle the smallest is Runtime: time limit: 49.7% of the target, but only 93 of the 3,550 rubrics (2.6%), so T = 93 ÷ 0.497 = 187 (186 after rounding each category's budget). Every other category is then cut to its share of 187: output artifact missing, unreproducible, or overwritten has 187 rubrics and keeps 3, parameter/threshold chosen without validation has 179 rubrics and keeps 1, fabricated, substituted, or partial input data has 131 rubrics and keeps 1.
share of MLE-bench Lite's failures (the target)share of the 3,550-rubric library
How big a matched library each category could supply (rubrics available ÷ the category's share of the target); the dashed line is the smallest, 187, which becomes the library's size. Only the ten smallest are shown.
All distributions side by side (% of each column)
category
collection pool errors
MLE-bench errors
collection pool, complete
MLE-bench, complete
wrong final answers
whole rollouts
weighted
oracle pool
oracle
weighted, 2nd+3rd
filtered oracle pool
filtered oracle
weighted, by weight
weighted, 2nd+3rd, by weight
Runtime: time limit
17.7
56.2
11.3
49.0
·
2.4
2.5
17.7
57.6
2.6
11.2
50.0
12.3
3.5
Runtime: exception in the code's own logic
33.5
15.6
23.7
13.6
·
6.2
5.3
33.6
15.8
4.9
24.1
14.0
16.4
7.2
Runtime: missing package
21.1
9.4
15.4
8.2
·
2.5
0.9
21.1
9.6
0.8
15.5
8.1
15.0
8.3
Missingness/validity unchecked before computation
15.2
7.0
12.5
6.1
6.0
21.8
16.1
15.3
7.3
15.5
12.3
6.5
11.1
7.2
Runtime: library API changed
7.5
6.2
4.8
5.4
·
0.8
1.4
7.5
6.2
1.4
4.8
5.4
2.9
1.9
Runtime: file not found
1.6
0.8
0.6
0.7
·
0.5
0.7
1.5
0.6
0.4
0.5
0.5
2.8
0.5
Output artifact missing, unreproducible, or overwritten
Rows/features altered before split without validation
0.3
·
3.0
·
8.2
4.4
5.1
0.2
·
5.4
3.2
·
0.7
2.1
Unsupervised clustering degeneracy unchecked
0.3
·
0.1
·
3.6
0.5
0.8
0.2
·
0.9
·
·
·
0.8
Runtime: tool misused
0.3
·
0.2
·
·
0.5
0.8
0.2
·
0.6
·
·
2.4
0.6
Output format/schema violates requirements
0.1
·
4.5
2.0
13.5
7.2
4.5
0.2
·
4.8
4.3
2.2
1.0
10.2
Numeric/identifier parsing or formatting corrupted
0.1
·
0.9
·
2.8
3.8
3.0
0.2
·
3.2
1.1
·
1.5
4.9
Runtime: out of memory
0.1
·
0.3
·
·
·
0.1
·
·
0.2
0.5
·
0.2
0.2
Model output not validated against baseline
·
·
1.6
10.2
12.5
11.4
16.3
·
·
15.5
1.6
10.2
5.5
11.9
Spec/config file ignored; rules invented
·
·
0.9
·
12.0
2.5
1.3
·
·
1.8
1.1
·
1.0
1.3
Statistic/aggregation computed on unverified scope
·
·
0.5
·
8.2
4.0
5.4
·
·
4.5
0.5
·
0.5
1.3
Data transformation without spec validation
·
·
0.1
·
0.6
1.4
2.5
·
·
2.1
·
·
0.6
1.1
Output/model inputs misaligned or unverified order
·
·
1.0
·
2.4
1.7
2.0
·
·
2.3
1.1
·
1.1
4.1
Derived quantity/extremum selection unchecked
·
·
0.0
·
0.3
1.1
1.6
·
·
1.5
·
·
·
0.5
Model optimized/thresholded to wrong metric
·
·
·
0.7
1.6
0.7
1.0
·
·
1.1
·
0.5
0.4
0.8
Entity aggregation uses only one role column
·
·
·
·
0.8
0.1
·
·
·
0.0
·
·
·
·
Visualization without persisted underlying data
·
·
·
·
0.3
0.3
0.4
·
·
0.5
·
·
·
0.2
Runtime: download failed
·
·
·
·
·
0.0
0.2
·
·
0.3
·
·
0.3
0.2
Script/output success
·
·
·
·
·
0.1
0.2
·
·
0.1
·
·
0.2
0.1
Modelling choice leaves score on the table
·
·
·
·
·
·
·
·
·
·
·
·
·
·
Runtime: dependency conflict
·
·
·
·
·
·
·
·
·
·
·
·
·
·
difference vs collection pool errors
0
42
28
47
91
69
71
1
42
71
29
47
38
68
difference vs MLE-bench errors
42
0
49
13
91
77
79
42
2
80
49
13
49
68
difference vs collection pool, complete
28
49
0
49
65
51
55
29
48
56
2
49
17
41
difference vs MLE-bench, complete
47
13
49
0
78
66
68
47
15
68
50
2
47
57
Hover a category for its definition. The collection pool's own errors are 42 away from MLE-bench's (49 on the complete targets): a library matched perfectly to the pool can get no closer to MLE-bench than about that.
Summary
library
rubrics
runtime share
vs collection pool, split / 21
vs MLE-bench, split / 21
vs pool, complete
vs MLE-bench, complete
collection pool errors (reference)
82%
0 / 0
42 / 10
28
47
MLE-bench errors (reference)
90%
42 / 10
0 / 0
49
13
collection pool, complete target (reference)
56%
28 / 28
49 / 34
0
49
MLE-bench, complete target (reference)
78%
47 / 14
13 / 13
49
0
Wrong final answers (baseline)
1,076
0%
91 / 91
91 / 91
65
78
Whole rollouts, no matching (baseline)
4,294
13%
69 / 69
77 / 77
51
66
Weighted
2,276
12%
71 / 70
79 / 78
55
68
Weighted, counted by weight
2,276
52%
38 / 34
49 / 38
17
47
Filtered oracle pool, first round (matched to the pool's failed commands)
464
82%
1 / 0
42 / 10
29
47
Filtered oracle, first round (matched to MLE-bench's failed commands)
177
90%
42 / 10
2 / 0
48
15
Weighted, second + third collections
3,550
11%
71 / 71
80 / 79
56
68
Weighted, second + third collections, counted by weight
3,550
22%
68 / 67
68 / 67
41
57
Filtered oracle pool (complete target)
187
57%
29 / 28
49 / 34
2
50
Filtered oracle (complete target)
186
78%
47 / 14
13 / 13
49
2
3. Which rubrics the judge was given, and the results
The judge never sees the whole library. Each time the agent runs code, the 12 rubrics most similar to the task and the code change are looked up (text embeddings; for the weighted library, the two ways described in its card), and the judge decides which apply and which are violated. Below: the library, the 12 given per call, and the ones the judge said were violated, against MLE-bench's errors. Every library here ran with the final harness (section 4); first-round libraries once (the weighted one twice per picking rule), second-round libraries twice.
Wrong final answers
MLE-bench errors (128)the librarythe 12 given to the judgejudged violated (147)
203 judge calls. Difference vs MLE-bench errors: library 91, given to the judge 98, judged violated 98. Runtime share of the 12 given: 0%.
Whole rollouts
MLE-bench errors (128)the librarythe 12 given to the judgejudged violated (147)
227 judge calls. Difference vs MLE-bench errors: library 77, given to the judge 96, judged violated 94. Runtime share of the 12 given: 1%.
Filtered oracle pool, first round
MLE-bench errors (128)the librarythe 12 given to the judgejudged violated (156)
272 judge calls. Difference vs MLE-bench errors: library 42, given to the judge 40, judged violated 46. Runtime share of the 12 given: 50%.
Filtered oracle, first round
MLE-bench errors (128)the librarythe 12 given to the judgejudged violated (73)
293 judge calls. Difference vs MLE-bench errors: library 2, given to the judge 23, judged violated 34. Runtime share of the 12 given: 79%.
Weighted (a), run 1
MLE-bench errors (128)the librarythe 12 given to the judgejudged violated (161)
262 judge calls. Difference vs MLE-bench errors: library 79, given to the judge 35, judged violated 44. Runtime share of the 12 given: 56%.
Weighted (a), run 2
MLE-bench errors (128)the librarythe 12 given to the judgejudged violated (166)
319 judge calls. Difference vs MLE-bench errors: library 79, given to the judge 38, judged violated 51. Runtime share of the 12 given: 52%.
Weighted (b), run 1
MLE-bench errors (128)the librarythe 12 given to the judgejudged violated (103)
213 judge calls. Difference vs MLE-bench errors: library 79, given to the judge 25, judged violated 34. Runtime share of the 12 given: 69%.
Weighted (b), run 2
MLE-bench errors (128)the librarythe 12 given to the judgejudged violated (152)
166 judge calls. Difference vs MLE-bench errors: library 79, given to the judge 28, judged violated 31. Runtime share of the 12 given: 69%.
2nd+3rd, 12 most similar
MLE-bench, complete target (147)the librarythe 12 given to the judgejudged violated (229)
220 judge calls. Difference vs MLE-bench, complete target: library 68, given to the judge 85, judged violated 86. Runtime share of the 12 given: 0%.
Weighted 2nd+3rd (a), run 1
MLE-bench, complete target (147)the librarythe 12 given to the judgejudged violated (212)
271 judge calls. Difference vs MLE-bench, complete target: library 68, given to the judge 63, judged violated 75. Runtime share of the 12 given: 18%.
Weighted 2nd+3rd (a), run 2
MLE-bench, complete target (147)the librarythe 12 given to the judgejudged violated (177)
271 judge calls. Difference vs MLE-bench, complete target: library 68, given to the judge 62, judged violated 74. Runtime share of the 12 given: 17%.
Filtered oracle pool, run 1
MLE-bench, complete target (147)the librarythe 12 given to the judgejudged violated (187)
237 judge calls. Difference vs MLE-bench, complete target: library 50, given to the judge 49, judged violated 45. Runtime share of the 12 given: 30%.
Filtered oracle pool, run 2
MLE-bench, complete target (147)the librarythe 12 given to the judgejudged violated (219)
277 judge calls. Difference vs MLE-bench, complete target: library 50, given to the judge 49, judged violated 60. Runtime share of the 12 given: 31%.
Filtered oracle, run 1
MLE-bench, complete target (147)the librarythe 12 given to the judgejudged violated (130)
229 judge calls. Difference vs MLE-bench, complete target: library 2, given to the judge 36, judged violated 38. Runtime share of the 12 given: 48%.
Filtered oracle, run 2
MLE-bench, complete target (147)the librarythe 12 given to the judgejudged violated (98)
261 judge calls. Difference vs MLE-bench, complete target: library 2, given to the judge 36, judged violated 43. Runtime share of the 12 given: 48%.
End to end on MLE-bench Lite (valid submissions / medals / above median, of 22)
library
how the 12 are chosen
harness
run 1
run 2
given to the judge: difference to MLE-bench, run 1 / 2
judged violated: difference, run 1 / 2
Reference: real execution, every command runs
–
no judge
18 / 2 / 3
–
–
Reference: no feedback, code-running commands do not run
–
no judge
7 / 2 / 2
–
–
Wrong final answers (baseline)
12 most similar
final (section 4)
12 / 2 / 2
98
98
Whole rollouts, no matching (baseline)
12 most similar
final (section 4)
11 / 2 / 2
96
94
Weighted
(a) split by kind in proportion to weight
final (section 4)
17 / 2 / 2
16 / 1 / 2
35 / 38
44 / 51
Weighted
(b) weight > 0 only, split by weight squared
final (section 4)
13 / 3 / 4 (19 of 22 finished)
14 / 2 / 2 (20 of 22 finished)
25 / 28
34 / 31
Filtered oracle pool, first round (matched to the pool's failed commands)
12 most similar
final (section 4)
17 / 0 / 3
40
46
Filtered oracle, first round (matched to MLE-bench's failed commands)
12 most similar
final (section 4)
16 / 0 / 1
23
34
Weighted, second + third collections
12 most similar, weight not used (no matching)
final (section 4)
14 / 2 / 2
85
86
Weighted, second + third collections
(a) split by kind in proportion to weight
final (section 4)
14 / 1 / 2
16 / 2 / 3
63 / 62
75 / 74
Filtered oracle pool (complete target)
12 most similar
final (section 4)
20 / 2 / 3
17 / 2 / 2
49 / 49
45 / 60
Filtered oracle (complete target)
12 most similar
final (section 4)
14 / 0 / 1
17 / 1 / 1
36 / 36
38 / 43
Difference columns: first-round rows against MLE-bench's trace failures, second-round rows against its complete target. Two identical runs of one configuration differ by 2-3 valid submissions and about 1 medal (the filtered oracle pool: 20 and 17), so single runs separate only large effects. First-round libraries ran once (2026-09-27), the second round twice (2026-09-28), all on non-Spot machines. The reference rows and the weighted runs are earlier runs of the same harness. Every run except the weighted (b) ones was checked for final runs that were cut off (status box) and those competitions re-run once; one in the second weighted 2nd+3rd run was cut off twice and counts as a failure. The weighted (b) runs are missing 3 + 2 competitions whose machines were reclaimed by Google mid-run or stalled while the laptop slept. text-normalization-challenge-english-language is invalid in every run including real execution: its answer file has 16 empty answers, which the grader loads as missing numbers and then cannot compare with text, so 21 is the reachable maximum. For reference, the library the final harness was developed with (rubrics written from each crashed command, 431) scores 16 / 2 / 2 and 17 / 2 / 2.
Reading. The two libraries with no matching and no weight (wrong final answers, whole rollouts) sit at 11-12 valid. Every library shaped toward the agent's failures sits at 14-20: oracle pool 17, oracle 16, weighted 17 / 16 in the first round; in the second, the filtered oracle pool 20 / 17, the weighted 2nd+3rd library 14 / 16 and the filtered oracle 14 / 17. The filtered oracle pool on the complete target is the only library whose two runs average at the 18 of running every command; the gap to the other shaped libraries (1-4) is inside run-to-run noise. Matching MLE-bench's own mix (the oracle, which looks at the test set) does no better than matching the pool, in both rounds. The outcome weight moved the weighted library's picks toward wrong-result rubrics (runtime share of what the judge got fell from 52-56% to 17-18%) and did not raise valid submissions. The 2nd+3rd library with no matching scores 14 against 11 for the first whole-rollout library, a gap at the noise level. Medals stay at 0-2 everywhere.
4. Harness fixes that improved results, independent of the library
These runs used the crashed-command library (431) and the Opus judge; each row adds one change to the row above. Valid submissions / medals / above median, of 22.
configuration
run 1
run 2
what changed
No feedback
7 / 2 / 2
Code-running commands do not run and return a fixed notice; no judge. Reference.
Judge only
10 / 3 / 3
8 / 2 / 2
Code-running commands do not run. The judge grades the code change against the 12 rubrics looked up for it; the agent sees a predicted reward and each applicable rubric with its own 0/1. The two basics listed below are already in.
+ the judge sees the sandbox
12 / 2 / 3
A read-only snapshot of the sandbox for the judge: Python version, installed packages, CPU/RAM/GPU, time left, file listings, refreshed after installs (CWM_WORLD_STATE). Before this, 50-64% of the scripts the judge passed crashed at an import or a file path on their first real run.
+ installs always run; package versions shown
8 / 2 / 2
Installs written as `cd X && pip install ... | tail` were mistaken for code execution and sent to the judge, so the packages were never installed (1-4 competitions per run). The install check now accepts cd, export and output filters, and the package list shows versions so a changed library API is visible.
+ really run the command when the judge finds nothing wrong
14 / 2 / 2
12 / 1 / 2
When no applicable rubric is violated, the command actually runs and the agent sees the real output (CWM_ESCALATE=green). The first change bigger than the difference between two identical runs.
+ data profile, last real run shown to the judge (3-competition test)
0 / 0 / 0 (3 of 22 finished)
The judge also sees each data file's shape, column types, missing values and odd file names (CWM_DATA_SAMPLE), and the last real run's command, exit code and output tail (CWM_REAL_MEMORY). It predicted all three targeted crashes, and all three still failed: the agent stopped running code and submitted despite a failing verdict.
+ the submission itself runs before it is accepted (final harness)
16 / 2 / 2
17 / 2 / 2
On submit the harness runs `python final.py` for real (capped) and grades it; the submission is refused, at most twice, while it crashes or a rubric is violated (CWM_SUBMIT_VERIFY=2). The run on a clean verdict also runs final.py.
Check: no feedback + the submission runs before it is accepted, no judge
10 / 1 / 2
10 / 1 / 2
Same submit check, judge never called: running the submission is worth about +3 valid over no feedback; the judge on top of it about +6-7.
Real execution: every command runs
18 / 2 / 3
The upper reference.
Two basics in every judge configuration
The submission script runs once after the episode (MLE_FINAL_SCRIPT): the agent writes its pipeline to final.py and the harness runs it once after the episode (30 min cap) to produce the submission; its output is saved as final_run.log, which is how the fixes above were found.
A fixed explanation at the top of each judge reply (CWM_FEEDBACK_PREAMBLE=1): the agent is told its code-running commands are not executed and what the reward and rubric lines mean.
PaperBench with the final harness: 0.230 mean score, against 0.194 with real execution and 0.198 with no feedback (paired difference vs no feedback +0.033, 95% CI [-0.001, +0.072]). MLE-bench Lite and PaperBench have both been used to find these failures, so every number on this page is a development-set number.