Rubric libraries: can a coding world model with rubrics stand in for running code?

Built 2026-10-06 16:21 · every library ever tried: detailed record · index

Experiment setupHaiku 4.5 agent; the coding world model (Opus 5 + a rubric library) predicts whether code works. Four feedback settings: Ceiling (all code runs), Floor (none, no feedback), Hybrid CWM (predictions plus a few capped real runs), Pure CWM (predictions only).
  • Agent: Claude Haiku 4.5 working in a sandbox on each task. Coding world model (CWM): Claude Opus 5 that reads the agent's code change against 12 rubrics from a library and predicts whether it works; the agent sees the predicted reward and each applicable rubric with its 0/1, never the CWM's reasoning.
  • Benchmarks: MLE-bench Lite (22 Kaggle competitions; official mle-bench grader: valid submission, medal, above median) and PaperBench (20 ICML 2024 papers to reproduce; score 0 to 1, graded by Claude Sonnet 5.5).
  • Where rubrics come from: the collection pool, the same Haiku agent's rollouts on other benchmarks (InfiAgent-DABench, DA-Code, DSBench, MLGym, MLAgentBench, RE-Bench for MLE-bench; SUPER research repositories for PaperBench), filtered so no task resembles a benchmark task. Every failure in it is labelled by Sonnet 5.5 into one of 22 categories (crashes split into their types). The category definitions were written once from failures of both the pool and MLE-bench's ceiling run (which kinds of failure exist, not how often); every proportion used to build or retrieve a library comes from the pool only.
  • Row names: curation method · retrieval method, and the feedback setting below.

What the agent sees, by feedback setting. Reading files, writing files and installing packages always happen in the sandbox; the settings differ only in what happens to commands that run code.

settinga command that runs codereal code runs during the episoderows
Ceilingexecuted; the agent sees the real outputevery code command (600s per-command limit)Ceiling
Floornot executed; a fixed notice, no feedbacknoneFloor
Hybrid CWMnot executed; the CWM predicts. When the prediction is clean, final.py runs for real (capped at 3 min) and the agent sees its real output; on submit, final.py runs for real (capped) before the submission is accepted; the CWM sees the last real outcomeclean-prediction checks and submit checks; 180s per check every rubric row of the MLE-bench Lite results table, the doubled-budget rows, full MLE-bench
Fixed-schedule checks
(control experiments only)
The harness counts the agent's code-running commands. The 1st, 3rd, 5th… are not executed: the agent sees the CWM prediction, or a fixed notice in the checks-only row. The 2nd, 4th, 6th… run final.py for real, whatever the command was, and the agent sees the real output (CWM rows also show the prediction). This is a command count, not a timer: 180s is how long each check may run before it is stopped. On submit, final.py runs again: a crash holds the submission (at most twice); reaching the 180s cap does not; the CWM cannot block a submission. Unlike Hybrid CWM, when checks happen never depends on the CWM's prediction. every second code command, plus submit; up to 180s eachthe fixed-policy control sections
Pure CWMnot executed; the CWM predicts; nothing elsenonethe pure world-model section

In every setting the harness runs final.py once after the episode ends, to produce the submission that is graded; the agent never sees that run. The CWM also sees a read-only view of the sandbox (installed packages, hardware, files) and a profile of the data, which inspect the environment, never the agent's code (a 2-second snapshot before each prediction, not counted as a code run).

Rubric curation methodsHow each library is built: Wrong-answer (the original), Crash-rubric, Unmatched, Filtered, Quota, and two Oracle references.
methodrubricsOpus cost per rubrichow the library is built
Wrong-answer1,000$0.05 ($46 total)The original: one rubric per rollout whose final answer was wrong, written from its final scripts and answer only.
Crash-rubric963 (PaperBench)$0.03 ($30 total)One rubric per crashed command of the collection pool. No matching.
Unmatched4,294$0.04 ($160 total)Up to 6 rubrics per rollout, written from the whole rollout (every command and output). Everything kept.
Filtered220$1.44 ($317 total)The whole-rollout library cut to the collection pool's failure proportions; the scarcest category sets the size.
Oracle187 (pool-matched)
186 (MLE-bench-matched)
$1.70 ($317 total)References, not methods: the Filtered recipe cut to the pool's exact counts (pool-matched) or to MLE-bench Lite's own failures (MLE-bench-matched, looks at the benchmark).
Quota952 (MLE-bench pool)
628 (PaperBench pool)
$0.07 ($70 total)Size and the pool's failure mix fixed first; rubrics written from each category's own failures until its quota is full.

Opus 5 mining calls only ($5 / $25 per million input / output tokens): a sample of 12 real prompts per miner re-issued to measure tokens (thinking included), times the calls each miner made, divided by the library's rubrics. Filtered and Oracle pay for the whole 3,550-rubric library they are cut from. Sonnet labelling of pool failures (needed by Filtered and Quota) not counted. MLE-bench libraries (analysis/mining_cost.py).

Retrieval methodsWhich 12 rubrics the CWM sees: Similarity (the 12 most similar) or Category split (12 slots divided by the collection pool's failure mix, then most similar within each category).
SimilarityAt each prediction the coding world model gets the 12 rubrics most similar to the task and the code change.
Category splitThe 12 slots are divided across failure categories in proportion to each category's share of the library's weight (weight = collection-pool failures a rubric was written from; 1 for Filtered), fixed per library. At each prediction each category's slots go to its most similar rubrics. Shares use pool-derived library support, not raw pool proportions or MLE-bench counts. Integer allocation can give a rare category zero slots on every request.

Category split, the slots each library gets:

categorycollection-pool shareQuota: slots of 12Filtered: slots of 12
Crash: exception in the code's own logic47%66
Crash: time limit13%22
Crash: missing package10%21
Output format/schema violates requirements6%11
Missingness/validity unchecked before computation4%11
Crash: library API changed3%01
The collection pool for MLE-bench Lite316 tasks, 2,795 rollouts from 6 benchmarks; InfiAgent-DABench is 69% of the tasks and 8% of the failed commands.

The Haiku agent's own rollouts on other data-science and ML benchmarks, none resembling an MLE-bench competition. Each task was run several times (seeds). The libraries are mined from the 2nd to 4th collections below.

69%17%7%316tasks62%15%18%2,795rollouts8%12%43%19%17%2,721failed commands
benchmarkkind of tasktasksrollouts2nd collection3rd collection4th collectionfailed commandswrong answers
InfiAgent-DABenchquestions about one CSV2171,727859868–221142
DA-Codedata wrangling, plotting, small ML tasks53422210212–331404
DSBenchKaggle-style modelling competitions214931581681671,174260
MLGymML research tasks (training, games)1386––8651862
MLAgentBenchML experimentation tasks1166––6647446
RE-BenchAI R&D engineering task11––130
Total3162,7951,2271,2483202,721914

Failed commands: found by exit code and error text. The 1st collection (5,397 rollouts: InfiAgent-DABench 4,337, DA-Code 1,060) was used only by the original Wrong-answer pipeline.

The original curation learned only from wrong final answersIt read only rollouts graded wrong, and only their last scripts and answer: none of the 1,192 failed commands in its source reached it. Failed commands are 80% of the pool's failures.

What "final" means. After a rollout ends, its submitted answer is graded. The original miner (Wrong-answer) kept only rollouts graded wrong and gave Opus four things: the task, the last version of each script the agent wrote, the submitted answer and the grader's verdict. It never saw the commands the agent ran or their outputs: no tracebacks, time limits or missing packages, none of the intermediate steps. Rollouts graded correct were skipped even when commands failed along the way (395 of the 1,192 failed commands in its source). So its rubrics describe mistakes visible in the final code and answer (wrong metric, split or format), not the failures that happen while code runs, which is what the CWM predicts at every step.

failed commands during the runwrong final answer or result
first collection (the original's source)47%53%collection pool (2nd to 4th collections)80%20%MLE-bench Lite, every command run88%12%what the original method learned from100%

Failure events. First collection: failed commands found by exit code and error text; wrong final answers. Collection pool and MLE-bench: failed commands, plus the cause of each wrong result found at grading.

Two rollouts behind Wrong-answer rubrics (click to open):

dacode-ml-regression-002: 4 failed commands in the trace, final answer wrong · rubric written: Validation split and feature set that don't mirror the actual prediction setting

Task: Predict electricity prices in test.csv based on your dataset on Spanish electricity and weather. Save your prediction results in result.csv, using the column name "price actual".

The full trace

1cd /home/agent && ls -la /home/data/ran
2cd /home/agent && head -50 /home/data/energy_dataset.csv && echo "---" && wc -l /home/data/energy_dataset.csvran
3cd /home/agent && head -10 /home/data/test.csv && echo "---" && wc -l /home/data/test.csvran
4cd /home/agent && head -10 /home/data/weather_features.csv && echo "---" && wc -l /home/data/weather_features. …ran
5write explore.py → python /home/agent/explore.pyran
6write train_model.py → python /home/agent/train_model.pyscript error: ValueError: Input X contains NaN.
7write train_model.py → python /home/agent/train_model.pyscript error: ValueError: Shape of passed values is (28051, 39), indices imply (28051, 41)
8write debug.py → python /home/agent/debug.pyran
9write train_model.py → python /home/agent/train_model.pyscript error: ValueError: Shape of passed values is (28051, 39), indices imply (28051, 41)
10write debug2.py → python /home/agent/debug2.pyran
11write debug3.py → python /home/agent/debug3.pyran
12write train_model.py → python /home/agent/train_model.pyapi drift: TypeError: HistGradientBoostingRegressor.__init__() got an unexpected keyword argument 'subsample'
13write train_model.py → python /home/agent/train_model.pyran
14cd /home/agent && head -20 result.csv && echo "---" && tail -10 result.csv && wc -l result.csvran
15cd /home/agent && python << 'EOF' …ran
16write answer.txtran
17echo COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT && cat /home/agent/answer.txtran

What the original method read

The final scripts (debug.py, debug2.py, debug3.py, explore.py, train_model.py), the final answer and the grader's verdict:

expected -> file result.csv: WRONG/MISSING

and nothing of the trace. The rubric it wrote:

### Validation split and feature set that don't mirror the actual prediction setting
- **Applies when**: `task` -- a model must produce predictions for a supplied test file whose rows come from an ordered (e.g., time-indexed) process, and the scripts choose which of the shared columns to use as features and how to hold out data for validation.
- **Pattern**: The attempt builds features by blanket-excluding columns (dropping some that exist in *both* train and test and are strongly related to the target) and then measures quality with a random/shuffled train-test split of ordered data, reporting the resulting optimistic error/R² as evidence the submission is good — without ever checking predi
dacode-ml-competition-006: 3 failed commands in the trace, final answer wrong · rubric written: Optimizing/validating with a metric other than the one the task specifies

Task: Design a method to predict wine quality using the provided dataset. Your task is to generate predictions for test.csv based on competition guidelines. Save your results in submission.csv following the format specified in sample_submission.csv.

The full trace

1cd /home/agent && ls -la /home/data/ran
2cd /home/agent && head -20 /home/data/train.csv && echo "---" && head -5 /home/data/test.csv && echo "---" && …ran
3write explore_data.py → python /home/agent/explore_data.pymissing package: ImportError: cannot import name 'quadratic_weighted_kappa' from 'sklearn.metrics' (/usr/local/lib/python3.12/site-packages/sklearn/metrics/__init__.py
4write explore_data.py → python /home/agent/explore_data.pyran
5write train_model.py → python /home/agent/train_model.pyran
6write train_optimized_model.py → python /home/agent/train_optimized_model.pyscript error: TypeError: only integer scalar arrays can be converted to a scalar index
7write train_optimized_model.py → python /home/agent/train_optimized_model.pyscript error: TypeError: only integer scalar arrays can be converted to a scalar index
8write train_optimized_model.py → python /home/agent/train_optimized_model.pyran
9cd /home/agent && head -20 submission.csv && echo "..." && tail -10 submission.csvran
10write verify_submission.py → python /home/agent/verify_submission.pyran
11write answer.txtran
12wc -l /home/agent/answer.txt && tail -20 /home/agent/answer.txt && echo "---" && head -1 /home/agent/answer.tx …ran
13cp /home/agent/submission.csv /home/agent/answer.txt && wc -l /home/agent/answer.txt && head -5 /home/agent/an …ran
14write SOLUTION_SUMMARY.mdran
15cd /home/agent && echo "Verifying answer.txt format..." && head -3 answer.txt && echo "..." && wc -l answer.tx …ran
16echo COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT && cat /home/agent/answer.txtran

What the original method read

The final scripts (explore_data.py, train_model.py, train_optimized_model.py, verify_submission.py), the final answer and the grader's verdict:

expected -> file submission.csv: WRONG/MISSING

and nothing of the trace. The rubric it wrote:

### Optimizing/validating with a metric other than the one the task specifies
- **Applies when**: `task` -- the task states an explicit scoring metric (e.g., a rank/agreement/ordinal or cost-weighted metric) and the scripts must choose a model, tune it, or decide when the answer is good enough.
- **Pattern**: The scripts never implement or compute the stated metric; they select the "best" model and ensemble weights using a generic proxy (accuracy, weighted F1, plain log-loss) and treat the target as unordered classes, so the reported/validated numbers say nothing about the actual grading score. Often compounded by a contaminated check (the final ensemble is fit on all training rows and then 
Rubrics from whole rollouts still missed the pool's failure mixCrashes are 76% of the pool's failures but 13% of the whole-rollout library: the generation never counts how often a failure happens.

Reading every command of every rollout fixed what the miner could see, not how often it wrote about each failure. Crashes are 76% of the pool's failures and 13% of the 3,550-rubric library. The miner writes up to 6 rubrics per rollout and merges near-duplicates: a crash seen hundreds of times collapses into a few rubrics, while every modelling lesson is new text and is kept. Nothing in the generation counts frequency: per pool failure, the library holds 21× more rubrics for other failures than for crashes.

collection-pool failuresWrong-answer libraryUnmatched, 2nd+3rd collections library
Crash: exception in the code's own logic47%6%Crash: time limit13%3%Crash: missing package10%1%Output format/schema violates requirements6%20%8%Missingness/validity unchecked before computation4%6%15%Crash: library API changed3%2%Modelling choice leaves score on the table3%1%5%Output artifact missing, unreproducible, or overwritte3%9%4%everything else11%65%56%

Distance to the pool's mix: Wrong-answer 78, whole rollouts 64 (0 = same mix, 100 = no overlap). Categories: Sonnet 5.5 reading each rubric.

Matching the pool's failure mixFiltered and Quota match the library to the pool: distance to the pool's mix from 78 (Wrong-answer) to 4 (Quota); the 12 rubrics shown then follow the pool's mix through Category split.

Two curation methods match the library to the collection pool's failure mix: Filtered cuts the whole-rollout library to the pool's proportions, Quota writes rubrics per category until each category's share is filled. Category split retrieval then gives each category its share of the 12 rubrics the CWM sees. Every library so far, by date built:

distance to the collection-pool mixdistance to the MLE-bench execution mix
02040608010009-09 Wrong-answer788909-21 Crash-rubric575509-23 Unmatched627309-24 Weighted637309-27 Filtered, first round324609-28 Unmatched, 2nd+3rd collections647309-28 Quota, first version455309-28 Oracle, pool-matched274509-28 Oracle, MLE-bench-matched501309-29 Filtered44509-29 Quota448

Each library that was run, against MLE-bench execution

Click a library. Execution = the failures MLE-bench Lite shows when every command runs (the ceiling runs).

Mismatch with execution: library 89, shown to the CWM 85 · library vs the pool it was matched to 78 · crash share: execution 83%, library 0%, shown 0%

MLE-bench failures seen during executioncollection-pool failures (the target)library (1,000 rubrics)the 12 shown to the CWM
Crash: time limit50%13%Crash: exception in the code's own logic15%47%Crash: missing package8%10%Modelling choice leaves score on the table7%3%1%11%Output format/schema violates requirements2%6%20%7%Crash: library API changed5%3%Missingness/validity unchecked before computation3%4%6%1%Model output not validated against baseline2%2%11%52%Output artifact missing, unreproducible, or overwritte3%9%7%Crash: out of memory1%1%Fabricated, substituted, or partial input data2%8%1%Crash: file not found1%1%Crash: disk full1%Rows/features altered before split without validation1%5%2%everything else5%41%19%

Mismatch with execution: library 73, shown to the CWM 86 · library vs the pool it was matched to 62 · crash share: execution 83%, library 16%, shown 1%

MLE-bench failures seen during executioncollection-pool failures (the target)library (4,294 rubrics)the 12 shown to the CWM
Crash: time limit50%13%2%Crash: exception in the code's own logic15%47%8%Crash: missing package8%10%2%Modelling choice leaves score on the table7%3%2%12%Output format/schema violates requirements2%6%11%12%Crash: library API changed5%3%1%Missingness/validity unchecked before computation3%4%19%Model output not validated against baseline2%2%8%57%Output artifact missing, unreproducible, or overwritte3%6%8%Crash: out of memory1%1%Fabricated, substituted, or partial input data2%3%1%Crash: file not found1%1%1%Crash: disk full1%Rows/features altered before split without validation1%2%everything else5%34%9%

Mismatch with execution: library 45, shown to the CWM 46 · library vs the pool it was matched to 4 · crash share: execution 83%, library 76%, shown 44%

MLE-bench failures seen during executioncollection-pool failures (the target)library (220 rubrics)the 12 shown to the CWM
Crash: time limit50%13%13%23%Crash: exception in the code's own logic15%47%45%18%Crash: missing package8%10%10%1%Modelling choice leaves score on the table7%3%3%3%Output format/schema violates requirements2%6%6%15%Crash: library API changed5%3%6%1%Missingness/validity unchecked before computation3%4%5%5%Model output not validated against baseline2%2%1%13%Output artifact missing, unreproducible, or overwritte3%3%11%Crash: out of memory1%1%1%2%Fabricated, substituted, or partial input data2%2%2%Crash: file not found1%1%Crash: disk full1%Rows/features altered before split without validation1%3%everything else5%4%3%

Mismatch with execution: library 48, shown to the CWM 60 · library vs the pool it was matched to 4 · crash share: execution 83%, library 75%, shown 26%

MLE-bench failures seen during executioncollection-pool failures (the target)library (952 rubrics)the 12 shown to the CWM
Crash: time limit50%13%11%12%Crash: exception in the code's own logic15%47%49%14%Crash: missing package8%10%9%Modelling choice leaves score on the table7%3%3%27%Output format/schema violates requirements2%6%6%11%Crash: library API changed5%3%4%Missingness/validity unchecked before computation3%4%5%2%Model output not validated against baseline2%2%2%12%Output artifact missing, unreproducible, or overwritte3%3%10%Crash: out of memory1%1%1%Fabricated, substituted, or partial input data2%2%5%Crash: file not found1%1%1%Crash: disk full1%Rows/features altered before split without validation1%1%everything else5%5%7%

Mismatch with execution: library 45, shown to the CWM 54 · library vs the pool it was matched to 27 · crash share: execution 83%, library 60%, shown 29%

MLE-bench failures seen during executioncollection-pool failures (the target)library (187 rubrics)the 12 shown to the CWM
Crash: time limit50%13%12%22%Crash: exception in the code's own logic15%47%24%5%Crash: missing package8%10%12%2%Modelling choice leaves score on the table7%3%3%10%Output format/schema violates requirements2%6%6%10%Crash: library API changed5%3%7%1%Missingness/validity unchecked before computation3%4%9%8%Model output not validated against baseline2%2%3%13%Output artifact missing, unreproducible, or overwritte3%10%17%Crash: out of memory1%1%1%Fabricated, substituted, or partial input data2%2%3%Crash: file not found1%1%3%Crash: disk full1%Rows/features altered before split without validation1%1%4%everything else5%9%6%

Mismatch with execution: library 13, shown to the CWM 44 · library vs the pool it was matched to 50 · crash share: execution 83%, library 78%, shown 46%

MLE-bench failures seen during executioncollection-pool failures (the target)library (186 rubrics)the 12 shown to the CWM
Crash: time limit50%13%51%43%Crash: exception in the code's own logic15%47%13%1%Crash: missing package8%10%6%1%Modelling choice leaves score on the table7%3%2%1%Output format/schema violates requirements2%6%3%4%Crash: library API changed5%3%8%Missingness/validity unchecked before computation3%4%4%3%Model output not validated against baseline2%2%7%38%Output artifact missing, unreproducible, or overwritte3%2%1%Crash: out of memory1%1%Fabricated, substituted, or partial input data2%1%1%Crash: file not found1%1%1%Crash: disk full1%Rows/features altered before split without validation1%everything else5%3%6%

Mismatch with execution: library 47, shown to the CWM 49 · library vs the pool it was matched to 2 · crash share: execution 83%, library 76%, shown 83%

MLE-bench failures seen during executioncollection-pool failures (the target)library (220 rubrics)the 12 shown to the CWM
Crash: time limit50%13%13%17%Crash: exception in the code's own logic15%47%47%50%Crash: missing package8%10%10%8%Modelling choice leaves score on the table7%3%3%Output format/schema violates requirements2%6%6%8%Crash: library API changed5%3%4%8%Missingness/validity unchecked before computation3%4%5%8%Model output not validated against baseline2%2%1%Output artifact missing, unreproducible, or overwritte3%3%Crash: out of memory1%1%1%Fabricated, substituted, or partial input data2%2%Crash: file not found1%1%Crash: disk full1%Rows/features altered before split without validation1%everything else5%4%

Mismatch with execution: library 48, shown to the CWM 54 · library vs the pool it was matched to 4 · crash share: execution 83%, library 75%, shown 83%

MLE-bench failures seen during executioncollection-pool failures (the target)library (952 rubrics)the 12 shown to the CWM
Crash: time limit50%13%11%17%Crash: exception in the code's own logic15%47%49%50%Crash: missing package8%10%9%17%Modelling choice leaves score on the table7%3%3%Output format/schema violates requirements2%6%6%8%Crash: library API changed5%3%4%Missingness/validity unchecked before computation3%4%5%8%Model output not validated against baseline2%2%2%Output artifact missing, unreproducible, or overwritte3%3%Crash: out of memory1%1%1%Fabricated, substituted, or partial input data2%2%Crash: file not found1%1%1%Crash: disk full1%Rows/features altered before split without validation1%everything else5%5%
Fresh quality lessons and clean attribution controlsStopped by user: 125/264 episodes complete; queued repeats cancelled.

Audit of the 29 modelling-quality rubrics: 14 came from score contrasts across eight tasks; 15 came from wrong-output cases. All had support 1. Incomplete script bundles and confounded changes prevented a reliable final-program check. This audit does not establish that all 952 rubrics are bad.

Fresh collection: 128/128 attempts recorded; 125 valid; 0 collection errors, not retried. Sixteen new OpenML tasks, eight attempts each; 12 tasks for generation and four for development. Exact scripts, scored replay and all attempts are retained.

New-only quality library: 1 accepted lessons. Curation + review API cost per accepted rubric: $2.87 (collection excluded). A lesson needs source-pair support plus positive discrimination on two other tasks, including development, with no reversed development prediction. No old rubrics are blended in.

Partial replication: user-requested early stop. 6 queued arms were cancelled. Already-active arms finish; all results and unresolved failures are retained. Counts below keep the original planned denominators, not a completed four-repeat comparison.

Gated screen: compare the frozen new library on the existing eight validation tasks, which never enter curation. Reuse the existing Quota predictions rather than rejudge them. At least 12 lessons, complete coverage, better ordering and rubric-specific discrimination are required. Previously inspected validation is not fresh held-out evidence.

feedbackrubricscomplete episodesvalidmedalabove medianreal code minepisode + final min
Quota + CWM95243/88
1 unresolved
partial; further repeats cancelled––
Task-only CWM038/88
6 unresolved
partial; further repeats cancelled––
Checks only044/88partial; further repeats cancelled––
Quality-only CWM (conditional)10/88not run: screen gate blocked––

The four-repeat plan was stopped early by the user. Every row uses the same fixed-schedule checks (defined in the settings table): a real final.py run, stopped after at most 180 seconds, on every second code-running command, plus a real run at submit that holds only a crashed submission. No CWM submit veto. The quality-only row runs only after the screen passes and controls complete.

MATS only: eight Spot collection workers, then at most 16 non-Spot evaluation workers; never overlapping. No selective retries; unresolved infrastructure losses stay visible. The check policy—not realized calls or runtime—is fixed. Lite remains development evidence.

Why the quality library stopped: saved-evidence diagnosisWrong-program citations and applicability gates; one collection-accounting bug fixed.

Quality-curation funnel: 15 proposed lessons → 7 passed grounding/format checks → 6 unique → 1 accepted.

  • Wrong-program citations: 8 proposals failed because 21 quotes were attributed to the wrong program. The saved requests contained the correct programs; this was generated evidence misattribution, not a parser bug.
  • Applicability: 3 rejected source contrasts marked the weaker program as violating the lesson but the stronger program as outside its scope. The frozen gate requires both to be applicable.
  • Quality review: 4 lessons failed the requirement that every comparison pass its quality audit. These checks varied with the comparison context. Rejection is not proof that every proposed lesson is false.

All 114/114 reviews were usable. Offline replay reproduced the frozen selection and library exactly. Criteria, judgments and accepted lessons are unchanged.

Collection accounting: among 6 flagged records audited, 1 recovered successfully but retained a stale error; 3 failed before an episode began and 2 lacked a resolved final collection. The collector fix preserves failed-read history and clears the active error only after completed collection. Historical records and their conservative status remain unchanged.

Next: use a separately versioned protocol for correctly attributed evidence and both-applicable near-misses. This diagnosis does not establish an efficiency or medal improvement.

Saved-evidence audit, 2026-10-04. No new model calls, benchmark episodes or acceptance changes. Rejection reasons overlap. Evidence: curation audit · collection audit.

Modelling-quality curation: prospective repairCuration: insufficient valid lessons. 0 accepted lessons.

Completed run: where candidates were lost

StageObserved result
Generation (16 responses)4 invalid JSON; 4 omitted required evidence.
Source audits (8)3 mis-nested outputs; 1 malformed JSON.
Program reviews (256)246 usable; 3 missing responses; 5 JSON/prose failures; 2 wrong citation hashes.
Accepted lessons0 / 8. Pool test not run.

One candidate was blocked only by the restart's failed-authentication review. Another had only output-format gaps. These are not proof that the lessons are false; neither is accepted retrospectively.

Repair: v3 implements native structured outputs and harness-bound citation hashes. See the v3 section for run status; support gates remain unchanged. No benchmark scaling yet.

Offline diagnosis: 281 saved calls replayed, no new model calls. Frozen outputs reproduced exactly.

Curation: insufficient valid lessons. 0 accepted lessons. This is a new protocol, not a regrading of the original one-lesson library.

IssueProspective change
Quotes assigned to the wrong programBind evidence to a program hash and exact source lines; extract quotes locally. Independent review still checks whether that evidence supports the lesson.
Risk avoided confused with outside scopeSeparate input preconditions from the defect. A relevant program can satisfy a lesson by avoiding its risk. Inapplicable and uncertain decisions never count as positive support.
Quality votes changed across unrelated comparisonsOne blinded source/quality audit per lesson, then separate applicability and violation judgments.
The same program received inconsistent judgmentsOne lesson and one program per review, reused only when the task, source and public context are identical. No counterpart program or private scores are shown.

Same collection data: 19 valid-program contrasts; 16 eligible for generation. The original 12 training / four development task split is retained. No MLE-bench or validation examples enter curation.

Unchanged support gates: source contrast plus two other positive tasks, at least one development positive, no development reversals, complete reviews and at least 12 accepted lessons before the CWM screen. Scores remain observational evidence, not isolated causal effects.

No collection or Lite runs start automatically. A separately invoked paid curation can proceed to the existing pool screen only if its library passes the gate. More accepted lessons, better medal rates and a rubric-specific speedup are all still unproven.

Evidence: offline plan and verification. The original protocol, calls, decisions and library are preserved.

Modelling-quality curation: structured outputsCuration: insufficient valid lessons. 4 accepted lessons.

Curation: insufficient valid lessons. 4 accepted lessons. These changes affect offline rubric curation, not the coding agent's hybrid CWM feedback or interpreter checkpoints.

ChangeWhat it does
Structured outputsThe generator, source auditor and program reviewer must return the required JSON fields. Native API schemas constrain the format; local checks still reject invalid, refused or truncated answers.
Harness-bound citationsModels select source IDs and line ranges. The harness checks the spans and attaches exact source hashes. Correct formatting is not proof of a correct lesson.
Narrower claimsUse the satisfying example and near-miss to scope a rule, not invent a blanket prohibition or claim a score difference proves causality.

Same 19 pool contrasts, 16 eligible for generation. Same source/cross-task/development support gates and 12-lesson screen minimum. V1/v2 evidence is frozen, not repaired.

SkillRefiner-inspired comparison design

SkillRefiner summarizes full traces, clusters successes and failures separately, proposes recurring lessons, and checks supporting evidence. Its output is an agent instruction skill, not CWM rubrics.

  • Compare full-trace, outcome-aware cluster curation with final-program-pair curation. Separate crashes from valid-but-low-scoring programs; count support across tasks.
  • Compare matched-budget Haiku and Sonnet collection. Our traces are Haiku; the rubric writer is already Opus and the reviewer Sonnet. Keep the evaluated agent fixed.
  • Check that quality rubrics are actually retrieved, then measure private pool scores and efficiency under the same real-check policy. More rubrics alone do not establish benefit.

The paper does not establish MLE-bench medal gains or stronger-teacher transfer. We retain fail-closed evidence validation rather than its implementation's fail-open verifier.

Offline implementation evidence. No new validity, medal or above-median result.

Next quality experimentscomplete with blocked stages · updated 2026-10-05 08:18 UTC; pool-only comparisons, unchanged acceptance gates.

Compare full-trace, outcome-aware curation with final-program contrasts; then compare lessons collected by Haiku and Sonnet. The fresh teacher collections use the same 16 training tasks and eight attempts per task, not benchmark examples.

Continues after the workstation ran low on disk space. Completed teacher collections and completed curation results are not repeated. Saved model responses are hash-checked and reused; interrupted requests remain in the record and replacements are separately accounted for. The original deadline and acceptance checks are unchanged. This is a continuation, not another independent replicate.

Overnight continuations and monitoring

The R3 Haiku trace source stopped at 326/327 complete chunks; the R3 Sonnet trace source stopped at 260/261. Each rejected one summary citing an unsupplied source. These are source-study boundaries, not claims of recovered coverage or accepted lessons.

The native program-contrast corpora differ: Haiku has 23 pairs (19 training, four development), Sonnet has 22 (19 training, three development). Their source audits discarded 100 and 107 unsupported observations, respectively; these are not missing chunks or rejected lessons.

Each authorized continuation reuses validated work and permits one replacement summary plus its audit, with source IDs constrained to its request. Rejected responses remain recorded; neither continuation is another independent repeat. Coverage and lesson counts appear only when observed.

The bounded overnight observer publishes progress and flags failures, stalled active stages and low disk. It does not restart experiments or change their gates. The intervention runner retains its cloud queue and owned-pod cleanup.

Overnight observer: 2 recorded alert(s); see the monitoring report.

Why the stronger-teacher final-program library is still small

Of 25 candidates, 21 failed the confounding check and 23 lacked a positive development-task contrast. These counts overlap. One lesson passed: selecting a validated cutoff for a recall-balancing metric. This is not measured agent benefit.

ComparisonProgressCuration API cost / accepted rubricNext gate
Haiku collectioncomplete; reused, not recollected; 128/128 attempts recordedNot a rubric libraryScored artifacts and provenance checked before curation
Sonnet collectioncomplete; reused, not recollected; 128/128 attempts recordedNot a rubric libraryScored artifacts and provenance checked before curation
Existing Haiku · Full traceinsufficient valid lessons; 1 accepted lessonsUnknown / incompleteScreen: gated; agent comparison: gated
Fresh Haiku · Final programsinsufficient valid lessons; 0 accepted lessonsNo accepted rubricsScreen: gated; agent comparison: gated
Fresh Haiku · Full traceincomplete trace coverage; 326/327 trace chunks; lesson generation not reachedNo rubric-cost resultScreen: gated; agent comparison: gated
Fresh Sonnet · Final programsinsufficient valid lessons; 1 accepted lessonsUnknown / incompleteScreen: gated; agent comparison: gated
Fresh Sonnet · Full traceincomplete trace coverage; 260/261 trace chunks; lesson generation not reachedNo rubric-cost resultScreen: gated; agent comparison: gated
Fresh Haiku · Full trace
Citation-contract continuation (R4)
insufficient valid lessons; 1 accepted lessons; 327/327 trace chunksUnknown / incompleteCuration only; no automatic agent or benchmark run
Fresh Sonnet · Full trace
Citation-contract continuation (R4)
insufficient valid lessons; 1 accepted lessons; 261/261 trace chunksUnknown / incompleteCuration only; no automatic agent or benchmark run

Same wall, turn, command and sandbox limits for both teachers. API spend is not matched; a dollar-based early stop is disabled for both because the Sonnet SDK tariff is unreliable. This does not test expansion to new task families. Per-rubric cost is curation-only, excludes teacher collection, and is shown only when accounting is complete.

What can run after curation?

At least 12 lessons must pass the unchanged source, cross-task and development checks. Then 16 CWM probes on the eight reserved pool tasks test quality discrimination and actual rubric exposure. Only a passing screen enables the Haiku-agent comparison: new library, Quota, task-only CWM, and checks only; 32 episodes per arm. The every-second-code-request interpreter policy and real-only submit acceptance stay fixed.

Quality-only libraries replace, never blend with, Quota. The same small-tabular pool is development evidence; no new validity or medal claim, and no automatic MLE-bench or PaperBench launch.

How full-trace evidence is used

Summaries cite recorded actions and outputs and receive a separate grounding check. Recurring observations are clustered by task-local outcome (cosine ≥0.82, at least two training tasks). Proposed lessons must still pass checks on valid weak/strong program contrasts. An invalid trace can suggest a lesson, but support on other valid programs does not prove its failure was corrected.

At most eight Spot workers on MATS. This continuation uses two API-only curation lanes and launches no collection. Agent comparisons remain gated. Original reports stay preserved; no automatic episode retries or benchmark launches.

Controlled lesson interventions10 candidate lessons across 8 pool tasks · complete with original unknowns.

Does following a proposed lesson improve the same program? Cases are selected using source checks only, not favorable development scores. These are candidate lessons, not accepted library rubrics.

Each scoped correction is independently reviewed, then compared with the original on the same pool data and limits. Three repeated executions per version; at most 60 executions for this case list. 60 program executions attempted; 27 matched pairs valid on both sides.

The first attempt lost a GCP Spot node. Its 32 attempted slots are never replayed: 29 valid results and three unknowns remain in the record. A bounded continuation runs only the 28 untouched slots, with the same programs, runtime, scoring and original deadline. This table combines both portions of the same 60-slot schedule, not new replicates.

Proposed lessonPool taskCorrection reviewCorrection typeScore change
Shipping a blended predictor that cross-validation never scoredopenml-training-1479approvedModel selection+0.0661
Validation scores a different predictor than the one that is deployedopenml-training-1487approvedModel selection+0.1914
In-sample score used as the only model quality signalopenml-training-4534approvedDiagnostic only+0.0000
Unbounded tree depth asserted as tuned, with no capacity evidence in the scriptopenml-training-4534approvedModel selection+0.0000
Capacity limits hard-coded outside the validated search spaceopenml-training-1471approvedModel selection+0.0514
Threshold tuned on folds scored by estimators already fit on all training rowsopenml-training-1487approvedModel selection+0.1914
Asserting tighter-than-default capacity limits without any evidence in the scriptopenml-training-1462approvedModel selection+0.0000
Positional feature matrices built from two frames without verifying column correspondenceopenml-training-1485approvedRisk guard only+0.0000
Final estimator rebuilt from duplicated hyperparameter literals instead of the validated objectopenml-training-3approvedRisk guard only+0.0000
Transformer fitted on the whole training set before cross-validated model comparisonopenml-training-1464approvedModel selection+0.0000

Score change is corrected minus original balanced accuracy, averaged over paired-valid repeats. Validity failures and missing executions remain in the report; repeats are not independent datasets. Correction types were assigned before execution, not inferred from scores. Diagnostic-only and risk-guard corrections may leave predictions unchanged even when they satisfy the lesson.

Within-source mechanism diagnostic, not a medal, transfer, or agent-speedup result. At most four MATS Spot workers; only pool data. No new teacher collection or benchmark run.

Request-format correction

Two setup requests were rejected before any model response. Both attempts remain recorded. A free provider preflight identified an incompatible nullable-enum encoding; both corrected schemas now pass that preflight. Cases, models, local acceptance gates and execution limits are unchanged.

Curation from measured correctionscomplete; 2 improved / 1 worse / 5 not applied

Start with the completed, measured corrections. Generate new lessons from the actual before/after programs, then check whether applying a lesson helps on another task. Fresh correction proposals use public task context and code—not benchmark errors or scores.

Model responses: 239 / 239 reserved requests; 0 failed, 0 unresolved.

StageStateCuration / coverageMeasured code outcomes
Curate from completed measured correctionscomplete1 pass the source audit; 0 pass static screening—
Try the lessons on other training taskscomplete0 approved / 2 planned cases0 valid / 0 attempted code runs (12 planned)
Measure fresh corrections: first batchcomplete12 approved / 12 planned cases72 valid / 72 attempted code runs (72 planned)
2 improved / 8 unchanged / 2 worse; 12 fully paired cases
Curate from the expanded evidencecomplete1 pass the source audit; 0 pass static screening—
Measure fresh corrections: second batchcomplete10 approved / 12 planned cases60 valid / 60 attempted code runs (72 planned)
2 improved / 6 unchanged / 2 worse; 10 fully paired cases
Recheck the expanded candidate setcomplete2 pass the source audit; 1 pass static screening—
Freeze the final transfer selectionfrozen8 cases frozen—
Try it on separate development taskscomplete3 approved / 8 planned cases18 valid / 18 attempted code runs (48 planned)
2 improved / 0 unchanged / 1 worse; 3 fully paired cases

1 lesson passed static screening; the 12-lesson library gate was not met. Static judgments are not measured transfer success.

Actual transfer: 2 improved, 1 worse, 5 not applied
LessonCollection-pool development taskDecisionBalanced-accuracy change
Select the probability cutoff on held-out predictions instead of using the classifier's default hard labelopenml-training-38declinednot run
Select the probability cutoff on held-out predictions instead of using the classifier's default hard labelopenml-training-50approved+3.73 pp
Select the probability cutoff on held-out predictions instead of using the classifier's default hard labelopenml-training-311approved-3.89 pp
Select the probability cutoff on held-out predictions instead of using the classifier's default hard labelopenml-training-333approved+1.79 pp
Score the exact shipped blend rule out-of-fold and allow fallback to the best validated single modelopenml-training-38declinednot run
Score the exact shipped blend rule out-of-fold and allow fallback to the best validated single modelopenml-training-50declinednot run
Score the exact shipped blend rule out-of-fold and allow fallback to the best validated single modelopenml-training-311declinednot run
Score the exact shipped blend rule out-of-fold and allow fallback to the best validated single modelopenml-training-333declinednot run

pp = percentage points. Three deterministic repeats per executed case, not three independent tasks. Declined cases have no execution result. This mixed result does not establish robust transfer or CWM-agent efficiency gains.

Does better public-data context filter ineffective changes?

The same frozen proposals were reviewed with public CSV statistics added. The new reviews were frozen before reading old execution results; no programs were rerun.

Measured effectCasesApproved beforeApproved with public profile
Improved444
Unchanged14147
Worse444
Not measured100

23 reviewed proposals; 1 excluded without a frozen review/code pair. One review per case, with reused baseline judgments: this is a context diagnostic, not a replicated improvement or a test of agent speed. Scope approval does not promise a score gain.

Can harmful edits supply useful lessons?

4 measured harmful changes, viewed in reverse: the edited program is weak and the original is strong. Original execution records and score directions stay unchanged.

4 lessons generated; 0 passed the unchanged source-quality audit. Audit rejections concerned unsupported causal claims and unresolved alternative explanations. No accepted library, cross-task test, or new code execution.

Separate outcome evidence from deployable rubric text

Same four training examples and outcome-blind auditor; only the generation instruction changes. Private outcome explanations must stay out of the served criterion. No quality gate is relaxed.

Writing instructionLessons generatedPass source audit
Original40
Explicit separation of private evidence and criterion53

One generation per example. A source-audit pass is not evidence of transfer or an accepted library.

Do the better-written lessons distinguish improvements on other tasks?

Blind static reviews on all four previously measured beneficial corrections, covering three other training tasks. No new program execution or development selection.

LessonFavors better programFavors worse programNo distinction / unknownSupporting distinct tasks
Do not replace a default cutoff with an unsmoothed argmax over every observed probability value004 / 00
Standardize each row when row-level scale dominates column-level differences004 / 00
Require a noise margin or resampling check before adopting a tuned decision threshold013 / 00

24 / 24 program reviews usable. Pair counts are not independent tasks. This tests static discrimination on saved evidence, not the benefit of teaching the lesson to an agent. The 12-lesson library gate is unchanged.

Does showing another program stabilize CWM review?

Same 24 program/lesson inputs. The new arms use identical instructions, with either the target repeated or its counterpart shown as an unscored, non-citable comparison. Scores and outcome order remain hidden; lessons and citation checks are unchanged.

The saved-evidence audit found a curation defect: two threshold lessons require threshold tuning to apply, yet also accept avoiding it. The source audit missed this inconsistency. The diagnostic measures review consistency, not a repaired library.

Review contextUsable reviewsMixed scopeFavors worse programFavors better programNo distinction / unknown
Single program (saved baseline)24 / 2461011 / 0
Repeat target (control)21 / 244108 / 3
Other program (comparison)20 / 243207 / 3

State: complete with unknown. 4 / 19 comparable target decisions differ between the two new arms; 5 targets lack a valid decision in at least one arm. Contrast columns count 12 lesson/correction combinations on just three tasks. Mixed scope means one program is applicable and the other is not; it is not automatically an error. Uncertain reviews remain separate.

Here, beneficial edits add the threshold tuning these lessons discourage. More reversals can expose a bad lesson rather than a bad reviewer. A favorable classification is not a success endpoint on this panel. The reused baseline also differs in framing and sampling; no agent, efficiency or medal gain is established. 48 new API requests maximum; no new code execution or cloud resources.

Rejected lessons do not end the whole study: the next planned source-expansion batch follows. The original quality checks and 12-lesson library gate stay unchanged. Expensive cross-task reviews run only after a source audit passes; source rejections stay in the denominator. Source-audited candidates in a diagnostic are not an accepted CWM library.

Comparison and safety limits

Each approved correction is compared with its unchanged original under the same execution cap, data, metric and seeds, with three repetitions. A case has six code runs: original and correction, each repeated three times. These repetitions are not independent tasks. Improved / unchanged / worse refers to balanced accuracy on fully paired cases; screened-out cases have no measured effect. Original interrupted results remain unknown; neither spent executions nor source programs are silently repeated. Final development selections are frozen before their execution scores are read; those scores cannot be used to rewrite or reselect lessons.

Deadline: 2026-10-05T17:33:00.144Z. At most four CPU Spot pods on MATS; 288 new program attempts, 300 seconds each. Owned pods are removed at completion or stop. This is small-tabular collection-pool development—not CWM-agent efficiency, medals, or a held-out benchmark result. No agent or benchmark arm is launched by this study.

Larger curation comparison: learn from counterexamplescomplete with unknown · terminal; 256 requests reserved, 256 returned · updated 2026-10-06 00:04 UTC

29 measured collection-pool contrasts on 11 training tasks. Each fold excludes one whole task from generation. All conditions share source anchors, generation repetitions, models and validation. No benchmark or old development outcomes enter the study.

Single example uses one source pair. Multiple examples adds the other training evidence. Counterexample-aware uses the same additional evidence but explicitly reconciles improvements, regressions and unchanged results, and requires consistent scope for satisfying alternatives.

Writing conditionGeneration calls finishedLessons generatedSource audit: pass / reject / unknownFavors better / worse on excluded tasksUnknown contrasts
Single example44 / 4484 / 3 / 10 / 08
Multiple examples44 / 44105 / 2 / 30 / 03
Counterexample-aware44 / 4472 / 3 / 20 / 03

All review decisions are frozen. Unchanged-score cases and distinct-task support are reported separately from favorable ordering; missing responses remain unknown.

132 matched generation calls planned; at most 2580 model requests for the frozen design. Deadline: 2026-10-06T07:50:50.693106+00:00. Four concurrent calls; no automatic retries, new code runs or cloud resources. Scientific rejection advances the remaining groups. This is task-excluded TRAIN cross-validation, not a fresh benchmark result, an agent efficiency gain or an accepted library.

Matched lesson value: does a CWM lesson improve the correction?Complete · 4 paired TRAIN tasks · balanced accuracy: unchanged 83.06%, no lesson 80.82%, lesson-guided 82.90%

Compare each unchanged original with a Haiku task-only correction (no lesson) and a matched Haiku lesson-guided correction, selecting other TRAIN collection-pool programs marked violates before either correction or its outcomes. One lesson per correction: this tests lesson content, not the 12-rubric retrieval policy or a full agent episode.

complete; phase complete; updated 2026-10-06 01:59 UTC; deadline 2026-10-06T08:41:48.675293+00:00.

Request compatibility: 8 original requests returned HTTP 400; they remain failed requests, not negative lesson results. The native-schema compatibility check was provider_accepted in 9.42 seconds. The continuation reuses that response and runs only previously unstarted units. One intermediate schema-invalid response is also preserved, not repaired. The current wire schema is the exact native schema; explicit field instructions and native local validation remain. No deadline or overall request cap was increased.

Execution handoff: the first service could not find the cloud CLI tools. The corrected service environment passed a read-only MATS preflight. Execution reuses the frozen cases and approved programs; the earlier stop record is preserved. New model requests in this handoff: 0.

Model schedule scope_review: 28 / 28 finished; 0 active model calls. Requests: 286 reserved, 278 returned.

Curation: 9 grounded candidates from 9 planned generation requests; source audit 4 passed / 5 rejected or unknown (of 9); independent source consistency 4 / 4; 4 distinct lessons selected.

Target applicability: violates 59; satisfies 28; not applicable 76; insufficient evidence 0; unknown 18 (of 181 reviews). Selected: 19 cases on 10 tasks.

ArmCorrection cases: approved / declined / rejected / unknownExecution slots: attempted / plannedSlots: valid / invalid / unknown
Unchanged originalnot a correction57 / 5757 / 0 / 0
Haiku, no lesson9 / 1 / 9 / 027 / 5724 / 3 / 30
Haiku, lesson-guided8 / 0 / 11 / 024 / 5724 / 0 / 33

Comparison coverage: both corrections were approved in 4 / 19 cases across 4 tasks. Only paired-valid executions can establish their score difference; approval is not a score gain.

Local patch-validation failures before scope review: no lesson: 1; lesson-guided: 8. These are unusable proposed edits, not measured score losses.

Main contrast, lesson-guided minus no lesson: mean balanced-accuracy difference +0.0208; mean execution-seconds difference +0.7268. Paired-valid coverage 12 / 57; 45 pairs without both valid scores. Execution state: complete.

Same-case unchanged-program baseline

ArmMean balanced accuracyMean program seconds
Unchanged original83.06%22.21
No lesson80.82%21.95
Lesson-guided82.90%22.68

Identical 4 cases across 4 tasks in all three rows; 12 dependent repetitions. A gain over no-lesson edits is not necessarily a gain over leaving the program unchanged.

Lesson-guided: 3 better / 1 worse cases than no lesson; 2 better / 2 worse than the unchanged original.

3 no-lesson timeout executions belong to 1 case(s) without an executed lesson-guided counterpart. They are not a paired validity advantage.

Measured writer defects (earlier study): 131 / 132 generations were schema-valid but yielded only 25 native candidates; 111 emitted lessons had empty applicability. Independent reviews: 82 / 99 usable, 17 unusable, including 13 truncated. Prospective repair uses explicit nonblank-field instructions, native local validation and a 6,000-token review cap; formatting repair is not evidence of better lessons.

Source-reviewed is not deployment-approved: native gates and the 12-lesson library gate remain unchanged. Effects condition on both arms valid; declined, rejected, missing and ambiguous outcomes are not zero scores. Three repeats preserve program seeds; repeats and cases sharing a task are dependent. TRAIN-only reused tasks, not a benchmark efficacy result, CWM-agent efficiency result or accepted library.

Envelope: 8 concurrent API requests / 4 GKE workers on MATS only; up to 24 cases × 3 arms × 3 repeats = 216 program attempts, 1,000 API requests and eight hours. Public JSON: plan · progress and results · writer diagnosis · request compatibility · correction coverage · matched final results.

Full-program correction: separate editing failures from lesson valueexecution complete; interrupted outcomes retained · 19 cases across 10 reused TRAIN tasks

Change under test: return one complete corrected program instead of substring replacements, in both the no-lesson and lesson-guided arms. Same frozen cases, lessons, models, output budgets, seeds and validation gates; the unchanged original remains the third arm. No selection of previous winners.

The previous study lost 8 / 19 guided proposals and 1 / 19 unguided proposals before scope review. This experiment first measures usable correction coverage, then score and runtime against both controls. A formatting fix is not evidence of better lessons.

Status: execution complete; interrupted outcomes retained; phase stopped; deadline 2026-10-06T15:14:05.190416+00:00.

Current model schedule: 29 / 29 finished; 0 active. Total requests: 69 reserved, 69 returned.

Execution-only recovery after infrastructure stops. 1 previously unstarted executions; the earlier 114 valid executions and 4 interrupted/unknown outcomes are retained. 0 new model requests. No attempted program is retried; the original deadline and scientific gates are unchanged.

1 earlier invalid execution is also retained, not retried.

Finish the sole unstarted slot after the disk-reserve stop; retain every previous outcome.

Saved-output audit: three file-reference rejections came from changed module docstrings containing existing paths, not new file access. Three other outputs do not parse as Python; one uses blocked dynamic class construction. Original rejections stay in place. These checks do not establish whether the rejected edits would improve predictions.

ArmCases: approved / declined / rejected / unknownExecutions: attempted / plannedValid / invalid / no observed outcome
Unchanged originalnot a correction57 / 5756 / 0 / 1
No lesson8 / 0 / 11 / 024 / 5721 / 1 / 35
Lesson-guided13 / 0 / 6 / 039 / 5738 / 0 / 19

No observed outcome includes unapproved and unstarted slots, as well as interrupted executions.

Local construction/validation failures, No lesson: introduced_data_or_file_reference: 1; python_parse_error: 1.

Local construction/validation failures, Lesson-guided: introduced_data_or_file_reference: 2; introduced_or_changed_external_access: 1; python_parse_error: 2.

Same-case comparison

ArmMean balanced accuracyMean program seconds
Unchanged original77.16%6.48
No lesson73.82%6.73
Lesson-guided77.88%10.30

Identical 5 paired-valid cases across 5 tasks; 15 dependent repetitions. Means average each case's repetitions first; repeated executions are not new tasks.

Partial-evidence comparison: interrupted outcomes remain unknown and are excluded only from paired-valid means, never from the all-case accounting above.

Stop reason: all_admitted_slots_exhausted_partial_evidence.

All selected cases stay in the denominators. Rejected, declined and missing results are not zero scores. TRAIN-only diagnosis, not a medal-rate or full-agent efficiency result. No library deployed and no acceptance gate relaxed. Limit: 76 model requests, 8 concurrent; 4 MATS CPU workers, 300 seconds per execution, 3 repetitions and an eight-hour deadline. Public JSON: plan · progress and results · saved-output diagnosis.

Results on MLE-bench Lite (hybrid CWM)Hybrid CWM. Valid submissions: Quota · Category split 85% (18.8 of 22), Ceiling 75%, Floor 44%. Last result 2026-10-01 03:36.
curation methodrubricsrunsvalidmedalabove medianvalid submissions per run (of 22)library vs executionshown vs executionlibrary vs pool
References (no CWM)
Ceiling–475% (16.5)8%18%18 / 16 / 16 / 16–––
Floor–444% (9.8)5%9%7 / 15 / 9 / 8–––
Hybrid CWM · Retrieval: Similarity (the 12 most similar rubrics from the whole library). Rows marked reference are the oracle libraries, cut to a known failure mix
Wrong-answer1,000458% (12.8)8%12%14 / 13 / 13 / 11898578
Unmatched4,294457% (12.5)8%9%11 / 13 / 13 / 13738662
Filtered220467% (14.8)5%10%12 / 16 / 15 / 1645464
Quota952473% (16.0)7%11%15 / 16 / 15 / 1848604
Oracle, pool-matched (reference)187483% (18.2)8%11%20 / 17 / 19 / 17455427
Oracle, MLE-bench-matched (reference)186468% (15.0)5%7%14 / 17 / 12 / 17134450
Hybrid CWM · Retrieval: Category split (12 slots divided across failure categories, then most similar within each)
Filtered220475% (16.5)7%12%18 / 18 / 16 / 1447492
Quota952485% (18.8)5%9%20 / 16 / 20 / 1948544

Rates are over all finished runs of 22 competitions each (mean valid submissions in brackets). Runs of the same setup vary a lot (the floor: 7, 15 and 9 valid), so gaps of a few points are noise. Mismatch columns: total variation distance on the 22 failure categories with crashes split by type, 0 = identical mix, 100 = no overlap. Execution = the failures MLE-bench Lite shows when every command runs (three ceiling runs: failed commands plus the cause of each wrong result); pool = the collection pool's failures, which the libraries were matched to. Hover a header for its definition.

Recorded time against the ceiling

rowrecorded window / competitionceiling / row timereal execution during episode
Ceiling55.2 min1.0×39.9 min
Floor21.6 min2.6×0.0 min
Wrong-answer · Similarity21.9 min2.5×6.1 min
Filtered · Similarity24.5 min2.3×6.7 min
Filtered · Category split21.5 min2.6×7.7 min
Quota · Similarity25.2 min2.2×5.8 min
Quota · Category split26.2 min2.1×10.0 min

Historical records lack exact episode boundaries. This window includes logged setup waits and the final collection run, but not unlogged provisioning, grading or earlier retry attempts; it is not pure agent time. Means include valid and invalid episodes. Real execution excludes installs and read-only harness probes. The hybrid changes execution policy as well as feedback; this ratio does not isolate the rubric contribution.

What does the work: rubrics, CWM predictions, real checks, or the combinationMean valid submissions of 22: Floor 9.8; Real run at submit only 12.0; Checks on every change 15.2; Pure CWM, no rubrics 8.2; Pure CWM 11.8; Hybrid CWM, no rubrics 16.5; Hybrid CWM 18.8; Ceiling 16.5
rowCWM predictslibrary rubricsreal code runs during the episodefeedback: CWM only / real outputreal run min / competitionrunsvalidmedalabove medianvalid per run (of 22)
Floornononone0 / 0.00.0444% (9.8)5%9%7 / 15 / 9 / 8
Real run at submit onlynonoat submit0 / 0.71.2455% (12.0)7%9%10 / 10 / 15 / 13
Checks on every changenonoevery code change (final.py, 3 min cap) and at submit0 / 12.918.7469% (15.2)8%14%15 / 14 / 15 / 17
Pure CWM, no rubricsyes (task text only)nonone5.4 / 0.00.0438% (8.2)6%9%8 / 8 / 7 / 10
Pure CWMyesQuota, Category splitnone8.4 / 0.00.0453% (11.8)7%11%11 / 11 / 13 / 12
Hybrid CWM, no rubricsyes (task text only)noon clean predictions and at submit4.0 / 4.36.2475% (16.5)6%11%15 / 14 / 19 / 18
Hybrid CWMyesQuota, Category spliton clean predictions and at submit5.7 / 6.910.0485% (18.8)5%9%20 / 16 / 20 / 19
Ceilingnonoevery code command, 600s per-command limit0 / every run39.9475% (16.5)8%18%18 / 16 / 16 / 16

Same 22 competitions, same agent. The real run at submit is capped at 3 min and only holds a submission whose final.py crashes or that the CWM still flags. 'No rubrics': the CWM judges the change against the task statement alone (no library). The earlier scheduled-check sweep was stopped after GKE transport errors entered agent feedback; its partial grades are not included here or used for causal inference. It also differed in CWM submit vetoes. The fixed-policy controls below remove that veto and add a task-only CWM control. Feedback counts cover messages shown to the agent, not silent accepted-submit checks. Real-run time excludes environment/data inspection, installs and final collection. 'Checks on every change' runs the deliverable on every routed code request, without a CWM call.

Earlier fixed-policy controls: collection-loss caveat264/264 episodes graded; not a clean speedup estimate.

Same execution policy, fixed-schedule checks (defined in the settings table): the agent's 2nd, 4th, 6th… code-running commands each trigger a real run of final.py, stopped after at most 180 seconds; the other code commands are not executed. This counts commands; it is not a timer. At submit final.py runs again; only a real crash holds the submission (at most twice), and the CWM has no submit veto. The two CWM rows also show the CWM's prediction on every code command. Same agent and limits; Quota uses the frozen 952-rubric library, with no new mining. Four full 22-competition runs per row.

feedbackgradedvalidmedalabove medianreal runsreal minepisode minepisode + final run minlost finals
Quota rubrics + CWM88/8878.4%8.0%13.6%6.07.714.721.60
Task-only CWM88/8864.8%6.8%13.6%6.08.615.123.86
Checks only, no CWM88/8855.7%2.3%11.4%7.610.917.426.05
  • Quota rubrics + CWM minus Task-only CWM: validity difference (pp): +13.6 [+1.1, +28.4]; real-time difference (min): -0.9 [-3.1, +1.1]; episode + final difference (min): -2.2 [-5.4, +1.5]
  • Validity sensitivity: +6.8 to +13.6 pp if lost finals had failed or succeeded. Adverse-endpoint 95% CI: [-3.4, +19.3] pp.
  • Quota rubrics + CWM minus Checks only, no CWM: validity difference (pp): +22.7 [+11.3, +35.2]; real-time difference (min): -3.2 [-5.2, -1.2]; episode + final difference (min): -4.3 [-7.2, -1.5]
  • Validity sensitivity: +17.0 to +22.7 pp if lost finals had failed or succeeded. Adverse-endpoint 95% CI: [+4.5, +30.7] pp.
  • Task-only CWM minus Checks only, no CWM: validity difference (pp): +9.1 [-3.4, +21.6]; real-time difference (min): -2.3 [-4.5, -0.1]; episode + final difference (min): -2.2 [-6.0, +1.4]
  • Validity sensitivity: +3.4 to +15.9 pp if lost finals had failed or succeeded. Adverse-endpoint 95% CI: [-10.2, +18.2] pp.

Collection-loss caveat: all episodes remain in the denominator. Lost finals are not proven code failures; the observed CIs are not clean causal estimates. Sensitivity changes only missing final validity, not runtime. Nine agent-loop OOM terminations per row remain real failures. Times cover the last saved attempt, including failed outcomes and CWM latency; final collection is included in the last time column, but earlier retries, provisioning and grading are not. CIs resample competitions, keeping four repeats together. Same check rule does not force identical realized check counts or duration.

What this tests: Quota versus task-only isolates adding the curated library under this policy; Quota versus checks-only tests the whole feedback channel. Higher validity at longer runtimes is a tradeoff, not an efficiency win. These controls do not by themselves prove adaptive triggering or distribution matching is the cause.

Reading: the package has a validity advantage over checks even if their five lost finals all succeeded. The extra benefit of curated rubrics over task-only CWM is less certain after collection-loss sensitivity. A rubric-specific speedup is not established; medals are 7 versus 6, and above-median outcomes are 12 versus 12.

Where modelling-quality feedback is lost29 performance-gap rubrics in the library; 0 of 11,088 logged retrieved exposures.

Quality lessons exist, but this category is not retrieved. S9P01 is the existing category “Modelling choice leaves score on the table”.

stageperformance-gap category / totalshare
Pool failure-event reference138 / 4,8282.9%
Curated library29 / 9523.0%
Library support used for slot allocation29 / 1,7471.7%
Allocated retrieval slots0 / 120.0%
Logged library retrieval exposures0 / 11,0880.0%

Its support gives 0.20 of 12 slots before integer allocation, then 0 afterwards. The fixed split spends ten slots on crashes, timeouts and missing packages, one on output format and one on missing-value checks. Across 924 saved verdicts, S9P01 was never retrieved, marked applicable or flagged. Baseline-validation, metric-selection and hyperparameter-validation categories also had zero retrieved exposures.

Logged verdicts are not all grade-wrapper attempts or cached feedback deliveries. Other categories and the task criterion can still contain modelling advice. This establishes a coverage gap, not that filling it will win medals; training limits, CWM accuracy and agent use of advice remain untested.

Previous validation collection — probe: 64/64 new OpenML attempts recorded; 64 valid. Eight new small classification tasks, eight real interpreter attempts each. These tasks are reserved for validation, not rubric mining.

The frozen CWM comparison used current retrieval, two protected quality slots, and the task criterion alone. The protected-slot screen did not clear its complete-coverage gate, so it did not launch a Lite arm. The new curation experiment above uses separate training tasks, not these validation examples.

The existing collection had no unused task-disjoint examples. OpenML tests small-tabular quality discrimination, not medal gains on large competitions. Its datasets are public. Final collection now uses a durable exit record; a MATS smoke passed and its pod was deleted. New retries preserve all attempts rather than overwriting their costs.

Pure world model: no real code runs, library curated from scratch at each sizeValid submissions of 22 — 8 rubrics: 7.8, 99 rubrics: 12.2, 500 rubrics: 8.2, 952 rubrics: 11.8, 1,802 rubrics: 12.0.
051015208995009521,802rubrics in the library, curated from scratch (log scale)valid submissions of 22
rowrubricsdistance to the pool's mixrunsvalidmedalabove medianvalid per run (of 22)
Ceiling475% (16.5)8%18%18 / 16 / 16 / 16
Floor444% (9.8)5%9%7 / 15 / 9 / 8
Quota · Category split, hybrid CWM (952)485% (18.8)5%9%20 / 16 / 20 / 19
Pure CWM, Quota planned 10824435% (7.8)2%8%6 / 8 / 7 / 10
Pure CWM, Quota planned 100994456% (12.2)8%12%12 / 12 / 12 / 13
Pure CWM, Quota planned 5005001438% (8.2)7%8%6 / 8 / 11 / 8
Pure CWM, Quota planned 1,0009524453% (11.8)7%11%11 / 11 / 13 / 12
Pure CWM, Quota planned 2,0001,8028455% (12.0)7%10%11 / 11 / 12 / 14

Pure CWM: no code runs for real during the episode (no run on a clean prediction, none at submit); every library is curated from scratch by the Quota miner at that size (planned size; a category whose new failures were already covered stops early, so large sizes can come out smaller), served by Category split. The 1,000 row is the method's own library (952).

Does a longer agent budget help?Valid submissions of 22: Ceiling 2 h 16.5; Ceiling 4 h 16.5; Quota · Category split, hybrid CWM 2 h 18.8; Quota · Category split, hybrid CWM 4 h 18.0
rowagent budgetrunsvalidmedalabove medianvalid per run (of 22)
Ceiling2 h, 250 steps475% (16.5)8%18%18 / 16 / 16 / 16
Ceiling4 h, 500 steps275% (16.5)7%16%17 / 16
Quota · Category split, hybrid CWM2 h, 250 steps485% (18.8)5%9%20 / 16 / 20 / 19
Quota · Category split, hybrid CWM4 h, 500 steps282% (18.0)5%11%17 / 19

Doubled wall clock, step and cost limits; the per-command time limit (10 min) is unchanged.

Full MLE-bench (75 competitions)Valid submissions: Ceiling 75%, Floor 49%, Quota · Category split, hybrid CWM 84%
all competitionsthe 22 of Litethe 53 outside Lite
rowvalidmedalabove medianvalidmedalabove medianvalidmedalabove median
Ceiling75% (69)3%10%82% (22)0%14%72% (47)4%9%
Floor49% (69)1%4%41% (22)5%5%53% (47)0%4%
Quota · Category split, hybrid CWM (952 rubrics)84% (69)3%6%82% (22)5%9%85% (47)2%4%

One run per row of the full 75-competition MLE-bench, same harness as Lite (CPU pods, 2 h agent budget), counted on the competitions graded in all three rows (in brackets). Left out: three whose mle-bench graders fail on pandas 3 (tgs-salt, tensorflow2-question-answering, vinbigdata) and three whose ceiling pods were evicted in every attempt because the agent unpacked the data past the pod's 10 GB disk (freesound-audio-tagging-2019, inaturalist-2019, iwildcam-2019; the floor and the CWM row had valid submissions on the first two, so leaving them out favours the ceiling). The 53 competitions outside Lite were never used for error analysis or any choice: the first clean test of the method.

Results on PaperBenchMean score of 20 papers, graded by Sonnet 5.5. Last result 2026-09-29 21:25.
rowrubricsrun 1run 2vs floor, paired (95% CI)papers better / worse
Ceiling–0.1041/20 graded; grader failed on 19-0.000 [-0.035, +0.032]7 / 12
Floor–0.1050/20 graded; grader failed on 20––
Crash-rubric · Similarity (MLE-bench pool)9630.147+0.043 [+0.010, +0.078]14 / 5
Quota · Category split (PaperBench's own pool)6280.115+0.011 [-0.009, +0.030]12 / 8
Quota · Category split (MLE-bench pool, transfer)9520/20 graded; grader failed on 5, 15 not run––

Mean score over the 20 papers of the all split (90 min agent, 1 h reproduction), every arm graded by Claude Sonnet 5.5 (earlier arms regraded from their stored submissions). A difference is real only when its interval excludes 0. The Crash-rubric library was picked for PaperBench because its failure mix was closest to PaperBench's own (distance 35 vs 56 for PaperBench's pool library): that row has looked at the benchmark.

What the latest results say

  • Crash-rubric · Similarity (a library from the MLE-bench pool) is the only row clearly above the floor.
  • Quota · Category split from PaperBench's own pool is inside noise; that library came out short and off its target (628 of 999 rubrics: that pool's API and environment failures repeat the same few lessons).
  • Ceiling = floor. PaperBench's score is half code development (graded by reading the code) and half execution and results; the second half is near zero for every row, because nothing reproduces a paper's results in 90 minutes. So execution feedback barely moves the score, while a CWM that reads the code against rubrics targets the half that is earned.

No mismatch chart here: PaperBench's collection pool and its benchmark were labelled in different category sets.

Both benchmarks have been used to find what goes wrong, so these are development-set numbers.