Built 2026-10-06 16:21 · every library ever tried: detailed record · index
What the agent sees, by feedback setting. Reading files, writing files and installing packages always happen in the sandbox; the settings differ only in what happens to commands that run code.
| setting | a command that runs code | real code runs during the episode | rows |
|---|---|---|---|
| Ceiling | executed; the agent sees the real output | every code command (600s per-command limit) | Ceiling |
| Floor | not executed; a fixed notice, no feedback | none | Floor |
| Hybrid CWM | not executed; the CWM predicts. When the prediction is clean, final.py runs for real
(capped at 3 min) and the agent sees its real output; on submit, final.py runs for real (capped) before the submission
is accepted; the CWM sees the last real outcome | clean-prediction checks and submit checks; 180s per check | every rubric row of the MLE-bench Lite results table, the doubled-budget rows, full MLE-bench |
| Fixed-schedule checks (control experiments only) | The harness counts the agent's code-running commands.
The 1st, 3rd, 5th… are not executed: the agent sees the CWM prediction, or a fixed notice in the checks-only row.
The 2nd, 4th, 6th… run final.py for real, whatever the command was, and the agent sees the real output
(CWM rows also show the prediction). This is a command count, not a timer: 180s is how long each check may run before it is stopped.
On submit, final.py runs again: a crash holds the submission (at most twice); reaching the 180s cap does not; the CWM
cannot block a submission. Unlike Hybrid CWM, when checks happen never depends on the CWM's prediction. |
every second code command, plus submit; up to 180s each | the fixed-policy control sections |
| Pure CWM | not executed; the CWM predicts; nothing else | none | the pure world-model section |
In every setting the harness runs final.py once after the episode ends, to produce the submission that is
graded; the agent never sees that run. The CWM also sees a read-only view of the sandbox (installed packages, hardware, files) and a
profile of the data, which inspect the environment, never the agent's code (a 2-second snapshot before each prediction, not counted
as a code run).
| method | rubrics | Opus cost per rubric | how the library is built |
|---|---|---|---|
| Wrong-answer | 1,000 | $0.05 ($46 total) | The original: one rubric per rollout whose final answer was wrong, written from its final scripts and answer only. |
| Crash-rubric | 963 (PaperBench) | $0.03 ($30 total) | One rubric per crashed command of the collection pool. No matching. |
| Unmatched | 4,294 | $0.04 ($160 total) | Up to 6 rubrics per rollout, written from the whole rollout (every command and output). Everything kept. |
| Filtered | 220 | $1.44 ($317 total) | The whole-rollout library cut to the collection pool's failure proportions; the scarcest category sets the size. |
| Oracle | 187 (pool-matched) 186 (MLE-bench-matched) | $1.70 ($317 total) | References, not methods: the Filtered recipe cut to the pool's exact counts (pool-matched) or to MLE-bench Lite's own failures (MLE-bench-matched, looks at the benchmark). |
| Quota | 952 (MLE-bench pool) 628 (PaperBench pool) | $0.07 ($70 total) | Size and the pool's failure mix fixed first; rubrics written from each category's own failures until its quota is full. |
Opus 5 mining calls only ($5 / $25 per million input / output tokens): a sample of 12 real prompts per miner re-issued to measure tokens (thinking included), times the calls each miner made, divided by the library's rubrics. Filtered and Oracle pay for the whole 3,550-rubric library they are cut from. Sonnet labelling of pool failures (needed by Filtered and Quota) not counted. MLE-bench libraries (analysis/mining_cost.py).
| Similarity | At each prediction the coding world model gets the 12 rubrics most similar to the task and the code change. |
| Category split | The 12 slots are divided across failure categories in proportion to each category's share of the library's weight (weight = collection-pool failures a rubric was written from; 1 for Filtered), fixed per library. At each prediction each category's slots go to its most similar rubrics. Shares use pool-derived library support, not raw pool proportions or MLE-bench counts. Integer allocation can give a rare category zero slots on every request. |
Category split, the slots each library gets:
| category | collection-pool share | Quota: slots of 12 | Filtered: slots of 12 |
|---|---|---|---|
| Crash: exception in the code's own logic | 47% | 6 | 6 |
| Crash: time limit | 13% | 2 | 2 |
| Crash: missing package | 10% | 2 | 1 |
| Output format/schema violates requirements | 6% | 1 | 1 |
| Missingness/validity unchecked before computation | 4% | 1 | 1 |
| Crash: library API changed | 3% | 0 | 1 |
The Haiku agent's own rollouts on other data-science and ML benchmarks, none resembling an MLE-bench competition. Each task was run several times (seeds). The libraries are mined from the 2nd to 4th collections below.
| benchmark | kind of task | tasks | rollouts | 2nd collection | 3rd collection | 4th collection | failed commands | wrong answers |
|---|---|---|---|---|---|---|---|---|
| InfiAgent-DABench | questions about one CSV | 217 | 1,727 | 859 | 868 | – | 221 | 142 |
| DA-Code | data wrangling, plotting, small ML tasks | 53 | 422 | 210 | 212 | – | 331 | 404 |
| DSBench | Kaggle-style modelling competitions | 21 | 493 | 158 | 168 | 167 | 1,174 | 260 |
| MLGym | ML research tasks (training, games) | 13 | 86 | – | – | 86 | 518 | 62 |
| MLAgentBench | ML experimentation tasks | 11 | 66 | – | – | 66 | 474 | 46 |
| RE-Bench | AI R&D engineering task | 1 | 1 | – | – | 1 | 3 | 0 |
| Total | 316 | 2,795 | 1,227 | 1,248 | 320 | 2,721 | 914 |
Failed commands: found by exit code and error text. The 1st collection (5,397 rollouts: InfiAgent-DABench 4,337, DA-Code 1,060) was used only by the original Wrong-answer pipeline.
What "final" means. After a rollout ends, its submitted answer is graded. The original miner (Wrong-answer) kept only rollouts graded wrong and gave Opus four things: the task, the last version of each script the agent wrote, the submitted answer and the grader's verdict. It never saw the commands the agent ran or their outputs: no tracebacks, time limits or missing packages, none of the intermediate steps. Rollouts graded correct were skipped even when commands failed along the way (395 of the 1,192 failed commands in its source). So its rubrics describe mistakes visible in the final code and answer (wrong metric, split or format), not the failures that happen while code runs, which is what the CWM predicts at every step.
Failure events. First collection: failed commands found by exit code and error text; wrong final answers. Collection pool and MLE-bench: failed commands, plus the cause of each wrong result found at grading.
Two rollouts behind Wrong-answer rubrics (click to open):
Task: Predict electricity prices in test.csv based on your dataset on Spanish electricity and weather. Save your prediction results in result.csv, using the column name "price actual".
| 1 | cd /home/agent && ls -la /home/data/ | ran |
| 2 | cd /home/agent && head -50 /home/data/energy_dataset.csv && echo "---" && wc -l /home/data/energy_dataset.csv | ran |
| 3 | cd /home/agent && head -10 /home/data/test.csv && echo "---" && wc -l /home/data/test.csv | ran |
| 4 | cd /home/agent && head -10 /home/data/weather_features.csv && echo "---" && wc -l /home/data/weather_features. … | ran |
| 5 | write explore.py → python /home/agent/explore.py | ran |
| 6 | write train_model.py → python /home/agent/train_model.py | script error: ValueError: Input X contains NaN. |
| 7 | write train_model.py → python /home/agent/train_model.py | script error: ValueError: Shape of passed values is (28051, 39), indices imply (28051, 41) |
| 8 | write debug.py → python /home/agent/debug.py | ran |
| 9 | write train_model.py → python /home/agent/train_model.py | script error: ValueError: Shape of passed values is (28051, 39), indices imply (28051, 41) |
| 10 | write debug2.py → python /home/agent/debug2.py | ran |
| 11 | write debug3.py → python /home/agent/debug3.py | ran |
| 12 | write train_model.py → python /home/agent/train_model.py | api drift: TypeError: HistGradientBoostingRegressor.__init__() got an unexpected keyword argument 'subsample' |
| 13 | write train_model.py → python /home/agent/train_model.py | ran |
| 14 | cd /home/agent && head -20 result.csv && echo "---" && tail -10 result.csv && wc -l result.csv | ran |
| 15 | cd /home/agent && python << 'EOF' … | ran |
| 16 | write answer.txt | ran |
| 17 | echo COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT && cat /home/agent/answer.txt | ran |
The final scripts (debug.py, debug2.py, debug3.py, explore.py, train_model.py), the final answer and the grader's verdict:
expected -> file result.csv: WRONG/MISSING
and nothing of the trace. The rubric it wrote:
### Validation split and feature set that don't mirror the actual prediction setting - **Applies when**: `task` -- a model must produce predictions for a supplied test file whose rows come from an ordered (e.g., time-indexed) process, and the scripts choose which of the shared columns to use as features and how to hold out data for validation. - **Pattern**: The attempt builds features by blanket-excluding columns (dropping some that exist in *both* train and test and are strongly related to the target) and then measures quality with a random/shuffled train-test split of ordered data, reporting the resulting optimistic error/R² as evidence the submission is good — without ever checking predi
Task: Design a method to predict wine quality using the provided dataset. Your task is to generate predictions for test.csv based on competition guidelines. Save your results in submission.csv following the format specified in sample_submission.csv.
| 1 | cd /home/agent && ls -la /home/data/ | ran |
| 2 | cd /home/agent && head -20 /home/data/train.csv && echo "---" && head -5 /home/data/test.csv && echo "---" && … | ran |
| 3 | write explore_data.py → python /home/agent/explore_data.py | missing package: ImportError: cannot import name 'quadratic_weighted_kappa' from 'sklearn.metrics' (/usr/local/lib/python3.12/site-packages/sklearn/metrics/__init__.py |
| 4 | write explore_data.py → python /home/agent/explore_data.py | ran |
| 5 | write train_model.py → python /home/agent/train_model.py | ran |
| 6 | write train_optimized_model.py → python /home/agent/train_optimized_model.py | script error: TypeError: only integer scalar arrays can be converted to a scalar index |
| 7 | write train_optimized_model.py → python /home/agent/train_optimized_model.py | script error: TypeError: only integer scalar arrays can be converted to a scalar index |
| 8 | write train_optimized_model.py → python /home/agent/train_optimized_model.py | ran |
| 9 | cd /home/agent && head -20 submission.csv && echo "..." && tail -10 submission.csv | ran |
| 10 | write verify_submission.py → python /home/agent/verify_submission.py | ran |
| 11 | write answer.txt | ran |
| 12 | wc -l /home/agent/answer.txt && tail -20 /home/agent/answer.txt && echo "---" && head -1 /home/agent/answer.tx … | ran |
| 13 | cp /home/agent/submission.csv /home/agent/answer.txt && wc -l /home/agent/answer.txt && head -5 /home/agent/an … | ran |
| 14 | write SOLUTION_SUMMARY.md | ran |
| 15 | cd /home/agent && echo "Verifying answer.txt format..." && head -3 answer.txt && echo "..." && wc -l answer.tx … | ran |
| 16 | echo COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT && cat /home/agent/answer.txt | ran |
The final scripts (explore_data.py, train_model.py, train_optimized_model.py, verify_submission.py), the final answer and the grader's verdict:
expected -> file submission.csv: WRONG/MISSING
and nothing of the trace. The rubric it wrote:
### Optimizing/validating with a metric other than the one the task specifies - **Applies when**: `task` -- the task states an explicit scoring metric (e.g., a rank/agreement/ordinal or cost-weighted metric) and the scripts must choose a model, tune it, or decide when the answer is good enough. - **Pattern**: The scripts never implement or compute the stated metric; they select the "best" model and ensemble weights using a generic proxy (accuracy, weighted F1, plain log-loss) and treat the target as unordered classes, so the reported/validated numbers say nothing about the actual grading score. Often compounded by a contaminated check (the final ensemble is fit on all training rows and then
Reading every command of every rollout fixed what the miner could see, not how often it wrote about each failure. Crashes are 76% of the pool's failures and 13% of the 3,550-rubric library. The miner writes up to 6 rubrics per rollout and merges near-duplicates: a crash seen hundreds of times collapses into a few rubrics, while every modelling lesson is new text and is kept. Nothing in the generation counts frequency: per pool failure, the library holds 21× more rubrics for other failures than for crashes.
Distance to the pool's mix: Wrong-answer 78, whole rollouts 64 (0 = same mix, 100 = no overlap). Categories: Sonnet 5.5 reading each rubric.
Two curation methods match the library to the collection pool's failure mix: Filtered cuts the whole-rollout library to the pool's proportions, Quota writes rubrics per category until each category's share is filled. Category split retrieval then gives each category its share of the 12 rubrics the CWM sees. Every library so far, by date built:
Click a library. Execution = the failures MLE-bench Lite shows when every command runs (the ceiling runs).
Mismatch with execution: library 89, shown to the CWM 85 · library vs the pool it was matched to 78 · crash share: execution 83%, library 0%, shown 0%
Mismatch with execution: library 73, shown to the CWM 86 · library vs the pool it was matched to 62 · crash share: execution 83%, library 16%, shown 1%
Mismatch with execution: library 45, shown to the CWM 46 · library vs the pool it was matched to 4 · crash share: execution 83%, library 76%, shown 44%
Mismatch with execution: library 48, shown to the CWM 60 · library vs the pool it was matched to 4 · crash share: execution 83%, library 75%, shown 26%
Mismatch with execution: library 45, shown to the CWM 54 · library vs the pool it was matched to 27 · crash share: execution 83%, library 60%, shown 29%
Mismatch with execution: library 13, shown to the CWM 44 · library vs the pool it was matched to 50 · crash share: execution 83%, library 78%, shown 46%
Mismatch with execution: library 47, shown to the CWM 49 · library vs the pool it was matched to 2 · crash share: execution 83%, library 76%, shown 83%
Mismatch with execution: library 48, shown to the CWM 54 · library vs the pool it was matched to 4 · crash share: execution 83%, library 75%, shown 83%
Audit of the 29 modelling-quality rubrics: 14 came from score contrasts across eight tasks; 15 came from wrong-output cases. All had support 1. Incomplete script bundles and confounded changes prevented a reliable final-program check. This audit does not establish that all 952 rubrics are bad.
Fresh collection: 128/128 attempts recorded; 125 valid; 0 collection errors, not retried. Sixteen new OpenML tasks, eight attempts each; 12 tasks for generation and four for development. Exact scripts, scored replay and all attempts are retained.
New-only quality library: 1 accepted lessons. Curation + review API cost per accepted rubric: $2.87 (collection excluded). A lesson needs source-pair support plus positive discrimination on two other tasks, including development, with no reversed development prediction. No old rubrics are blended in.
Partial replication: user-requested early stop. 6 queued arms were cancelled. Already-active arms finish; all results and unresolved failures are retained. Counts below keep the original planned denominators, not a completed four-repeat comparison.
Gated screen: compare the frozen new library on the existing eight validation tasks, which never enter curation. Reuse the existing Quota predictions rather than rejudge them. At least 12 lessons, complete coverage, better ordering and rubric-specific discrimination are required. Previously inspected validation is not fresh held-out evidence.
| feedback | rubrics | complete episodes | valid | medal | above median | real code min | episode + final min |
|---|---|---|---|---|---|---|---|
| Quota + CWM | 952 | 43/88 1 unresolved | partial; further repeats cancelled | – | – | ||
| Task-only CWM | 0 | 38/88 6 unresolved | partial; further repeats cancelled | – | – | ||
| Checks only | 0 | 44/88 | partial; further repeats cancelled | – | – | ||
| Quality-only CWM (conditional) | 1 | 0/88 | not run: screen gate blocked | – | – | ||
The four-repeat plan was stopped early by the user. Every row uses the same fixed-schedule checks (defined in the settings table): a real final.py run, stopped after at most 180 seconds, on every second code-running command, plus a real run at submit that holds only a crashed submission. No CWM submit veto. The quality-only row runs only after the screen passes and controls complete.
MATS only: eight Spot collection workers, then at most 16 non-Spot evaluation workers; never overlapping. No selective retries; unresolved infrastructure losses stay visible. The check policy—not realized calls or runtime—is fixed. Lite remains development evidence.
Quality-curation funnel: 15 proposed lessons → 7 passed grounding/format checks → 6 unique → 1 accepted.
All 114/114 reviews were usable. Offline replay reproduced the frozen selection and library exactly. Criteria, judgments and accepted lessons are unchanged.
Collection accounting: among 6 flagged records audited, 1 recovered successfully but retained a stale error; 3 failed before an episode began and 2 lacked a resolved final collection. The collector fix preserves failed-read history and clears the active error only after completed collection. Historical records and their conservative status remain unchanged.
Next: use a separately versioned protocol for correctly attributed evidence and both-applicable near-misses. This diagnosis does not establish an efficiency or medal improvement.
Saved-evidence audit, 2026-10-04. No new model calls, benchmark episodes or acceptance changes. Rejection reasons overlap. Evidence: curation audit · collection audit.
| Stage | Observed result |
|---|---|
| Generation (16 responses) | 4 invalid JSON; 4 omitted required evidence. |
| Source audits (8) | 3 mis-nested outputs; 1 malformed JSON. |
| Program reviews (256) | 246 usable; 3 missing responses; 5 JSON/prose failures; 2 wrong citation hashes. |
| Accepted lessons | 0 / 8. Pool test not run. |
One candidate was blocked only by the restart's failed-authentication review. Another had only output-format gaps. These are not proof that the lessons are false; neither is accepted retrospectively.
Repair: v3 implements native structured outputs and harness-bound citation hashes. See the v3 section for run status; support gates remain unchanged. No benchmark scaling yet.
Offline diagnosis: 281 saved calls replayed, no new model calls. Frozen outputs reproduced exactly.
Curation: insufficient valid lessons. 0 accepted lessons. This is a new protocol, not a regrading of the original one-lesson library.
| Issue | Prospective change |
|---|---|
| Quotes assigned to the wrong program | Bind evidence to a program hash and exact source lines; extract quotes locally. Independent review still checks whether that evidence supports the lesson. |
| Risk avoided confused with outside scope | Separate input preconditions from the defect. A relevant program can satisfy a lesson by avoiding its risk. Inapplicable and uncertain decisions never count as positive support. |
| Quality votes changed across unrelated comparisons | One blinded source/quality audit per lesson, then separate applicability and violation judgments. |
| The same program received inconsistent judgments | One lesson and one program per review, reused only when the task, source and public context are identical. No counterpart program or private scores are shown. |
Same collection data: 19 valid-program contrasts; 16 eligible for generation. The original 12 training / four development task split is retained. No MLE-bench or validation examples enter curation.
Unchanged support gates: source contrast plus two other positive tasks, at least one development positive, no development reversals, complete reviews and at least 12 accepted lessons before the CWM screen. Scores remain observational evidence, not isolated causal effects.
No collection or Lite runs start automatically. A separately invoked paid curation can proceed to the existing pool screen only if its library passes the gate. More accepted lessons, better medal rates and a rubric-specific speedup are all still unproven.
Evidence: offline plan and verification. The original protocol, calls, decisions and library are preserved.
Curation: insufficient valid lessons. 4 accepted lessons. These changes affect offline rubric curation, not the coding agent's hybrid CWM feedback or interpreter checkpoints.
| Change | What it does |
|---|---|
| Structured outputs | The generator, source auditor and program reviewer must return the required JSON fields. Native API schemas constrain the format; local checks still reject invalid, refused or truncated answers. |
| Harness-bound citations | Models select source IDs and line ranges. The harness checks the spans and attaches exact source hashes. Correct formatting is not proof of a correct lesson. |
| Narrower claims | Use the satisfying example and near-miss to scope a rule, not invent a blanket prohibition or claim a score difference proves causality. |
Same 19 pool contrasts, 16 eligible for generation. Same source/cross-task/development support gates and 12-lesson screen minimum. V1/v2 evidence is frozen, not repaired.
SkillRefiner summarizes full traces, clusters successes and failures separately, proposes recurring lessons, and checks supporting evidence. Its output is an agent instruction skill, not CWM rubrics.
The paper does not establish MLE-bench medal gains or stronger-teacher transfer. We retain fail-closed evidence validation rather than its implementation's fail-open verifier.
Offline implementation evidence. No new validity, medal or above-median result.
Compare full-trace, outcome-aware curation with final-program contrasts; then compare lessons collected by Haiku and Sonnet. The fresh teacher collections use the same 16 training tasks and eight attempts per task, not benchmark examples.
Continues after the workstation ran low on disk space. Completed teacher collections and completed curation results are not repeated. Saved model responses are hash-checked and reused; interrupted requests remain in the record and replacements are separately accounted for. The original deadline and acceptance checks are unchanged. This is a continuation, not another independent replicate.
The R3 Haiku trace source stopped at 326/327 complete chunks; the R3 Sonnet trace source stopped at 260/261. Each rejected one summary citing an unsupplied source. These are source-study boundaries, not claims of recovered coverage or accepted lessons.
The native program-contrast corpora differ: Haiku has 23 pairs (19 training, four development), Sonnet has 22 (19 training, three development). Their source audits discarded 100 and 107 unsupported observations, respectively; these are not missing chunks or rejected lessons.
Each authorized continuation reuses validated work and permits one replacement summary plus its audit, with source IDs constrained to its request. Rejected responses remain recorded; neither continuation is another independent repeat. Coverage and lesson counts appear only when observed.
The bounded overnight observer publishes progress and flags failures, stalled active stages and low disk. It does not restart experiments or change their gates. The intervention runner retains its cloud queue and owned-pod cleanup.
Overnight observer: 2 recorded alert(s); see the monitoring report.
Of 25 candidates, 21 failed the confounding check and 23 lacked a positive development-task contrast. These counts overlap. One lesson passed: selecting a validated cutoff for a recall-balancing metric. This is not measured agent benefit.
| Comparison | Progress | Curation API cost / accepted rubric | Next gate |
|---|---|---|---|
| Haiku collection | complete; reused, not recollected; 128/128 attempts recorded | Not a rubric library | Scored artifacts and provenance checked before curation |
| Sonnet collection | complete; reused, not recollected; 128/128 attempts recorded | Not a rubric library | Scored artifacts and provenance checked before curation |
| Existing Haiku · Full trace | insufficient valid lessons; 1 accepted lessons | Unknown / incomplete | Screen: gated; agent comparison: gated |
| Fresh Haiku · Final programs | insufficient valid lessons; 0 accepted lessons | No accepted rubrics | Screen: gated; agent comparison: gated |
| Fresh Haiku · Full trace | incomplete trace coverage; 326/327 trace chunks; lesson generation not reached | No rubric-cost result | Screen: gated; agent comparison: gated |
| Fresh Sonnet · Final programs | insufficient valid lessons; 1 accepted lessons | Unknown / incomplete | Screen: gated; agent comparison: gated |
| Fresh Sonnet · Full trace | incomplete trace coverage; 260/261 trace chunks; lesson generation not reached | No rubric-cost result | Screen: gated; agent comparison: gated |
| Fresh Haiku · Full trace Citation-contract continuation (R4) | insufficient valid lessons; 1 accepted lessons; 327/327 trace chunks | Unknown / incomplete | Curation only; no automatic agent or benchmark run |
| Fresh Sonnet · Full trace Citation-contract continuation (R4) | insufficient valid lessons; 1 accepted lessons; 261/261 trace chunks | Unknown / incomplete | Curation only; no automatic agent or benchmark run |
Same wall, turn, command and sandbox limits for both teachers. API spend is not matched; a dollar-based early stop is disabled for both because the Sonnet SDK tariff is unreliable. This does not test expansion to new task families. Per-rubric cost is curation-only, excludes teacher collection, and is shown only when accounting is complete.
At least 12 lessons must pass the unchanged source, cross-task and development checks. Then 16 CWM probes on the eight reserved pool tasks test quality discrimination and actual rubric exposure. Only a passing screen enables the Haiku-agent comparison: new library, Quota, task-only CWM, and checks only; 32 episodes per arm. The every-second-code-request interpreter policy and real-only submit acceptance stay fixed.
Quality-only libraries replace, never blend with, Quota. The same small-tabular pool is development evidence; no new validity or medal claim, and no automatic MLE-bench or PaperBench launch.
Summaries cite recorded actions and outputs and receive a separate grounding check. Recurring observations are clustered by task-local outcome (cosine ≥0.82, at least two training tasks). Proposed lessons must still pass checks on valid weak/strong program contrasts. An invalid trace can suggest a lesson, but support on other valid programs does not prove its failure was corrected.
At most eight Spot workers on MATS. This continuation uses two API-only curation lanes and launches no collection. Agent comparisons remain gated. Original reports stay preserved; no automatic episode retries or benchmark launches.
Does following a proposed lesson improve the same program? Cases are selected using source checks only, not favorable development scores. These are candidate lessons, not accepted library rubrics.
Each scoped correction is independently reviewed, then compared with the original on the same pool data and limits. Three repeated executions per version; at most 60 executions for this case list. 60 program executions attempted; 27 matched pairs valid on both sides.
The first attempt lost a GCP Spot node. Its 32 attempted slots are never replayed: 29 valid results and three unknowns remain in the record. A bounded continuation runs only the 28 untouched slots, with the same programs, runtime, scoring and original deadline. This table combines both portions of the same 60-slot schedule, not new replicates.
| Proposed lesson | Pool task | Correction review | Correction type | Score change |
|---|---|---|---|---|
| Shipping a blended predictor that cross-validation never scored | openml-training-1479 | approved | Model selection | +0.0661 |
| Validation scores a different predictor than the one that is deployed | openml-training-1487 | approved | Model selection | +0.1914 |
| In-sample score used as the only model quality signal | openml-training-4534 | approved | Diagnostic only | +0.0000 |
| Unbounded tree depth asserted as tuned, with no capacity evidence in the script | openml-training-4534 | approved | Model selection | +0.0000 |
| Capacity limits hard-coded outside the validated search space | openml-training-1471 | approved | Model selection | +0.0514 |
| Threshold tuned on folds scored by estimators already fit on all training rows | openml-training-1487 | approved | Model selection | +0.1914 |
| Asserting tighter-than-default capacity limits without any evidence in the script | openml-training-1462 | approved | Model selection | +0.0000 |
| Positional feature matrices built from two frames without verifying column correspondence | openml-training-1485 | approved | Risk guard only | +0.0000 |
| Final estimator rebuilt from duplicated hyperparameter literals instead of the validated object | openml-training-3 | approved | Risk guard only | +0.0000 |
| Transformer fitted on the whole training set before cross-validated model comparison | openml-training-1464 | approved | Model selection | +0.0000 |
Score change is corrected minus original balanced accuracy, averaged over paired-valid repeats. Validity failures and missing executions remain in the report; repeats are not independent datasets. Correction types were assigned before execution, not inferred from scores. Diagnostic-only and risk-guard corrections may leave predictions unchanged even when they satisfy the lesson.
Within-source mechanism diagnostic, not a medal, transfer, or agent-speedup result. At most four MATS Spot workers; only pool data. No new teacher collection or benchmark run.
Two setup requests were rejected before any model response. Both attempts remain recorded. A free provider preflight identified an incompatible nullable-enum encoding; both corrected schemas now pass that preflight. Cases, models, local acceptance gates and execution limits are unchanged.
Start with the completed, measured corrections. Generate new lessons from the actual before/after programs, then check whether applying a lesson helps on another task. Fresh correction proposals use public task context and code—not benchmark errors or scores.
Model responses: 239 / 239 reserved requests; 0 failed, 0 unresolved.
| Stage | State | Curation / coverage | Measured code outcomes |
|---|---|---|---|
| Curate from completed measured corrections | complete | 1 pass the source audit; 0 pass static screening | — |
| Try the lessons on other training tasks | complete | 0 approved / 2 planned cases | 0 valid / 0 attempted code runs (12 planned) |
| Measure fresh corrections: first batch | complete | 12 approved / 12 planned cases | 72 valid / 72 attempted code runs (72 planned) 2 improved / 8 unchanged / 2 worse; 12 fully paired cases |
| Curate from the expanded evidence | complete | 1 pass the source audit; 0 pass static screening | — |
| Measure fresh corrections: second batch | complete | 10 approved / 12 planned cases | 60 valid / 60 attempted code runs (72 planned) 2 improved / 6 unchanged / 2 worse; 10 fully paired cases |
| Recheck the expanded candidate set | complete | 2 pass the source audit; 1 pass static screening | — |
| Freeze the final transfer selection | frozen | 8 cases frozen | — |
| Try it on separate development tasks | complete | 3 approved / 8 planned cases | 18 valid / 18 attempted code runs (48 planned) 2 improved / 0 unchanged / 1 worse; 3 fully paired cases |
1 lesson passed static screening; the 12-lesson library gate was not met. Static judgments are not measured transfer success.
| Lesson | Collection-pool development task | Decision | Balanced-accuracy change |
|---|---|---|---|
| Select the probability cutoff on held-out predictions instead of using the classifier's default hard label | openml-training-38 | declined | not run |
| Select the probability cutoff on held-out predictions instead of using the classifier's default hard label | openml-training-50 | approved | +3.73 pp |
| Select the probability cutoff on held-out predictions instead of using the classifier's default hard label | openml-training-311 | approved | -3.89 pp |
| Select the probability cutoff on held-out predictions instead of using the classifier's default hard label | openml-training-333 | approved | +1.79 pp |
| Score the exact shipped blend rule out-of-fold and allow fallback to the best validated single model | openml-training-38 | declined | not run |
| Score the exact shipped blend rule out-of-fold and allow fallback to the best validated single model | openml-training-50 | declined | not run |
| Score the exact shipped blend rule out-of-fold and allow fallback to the best validated single model | openml-training-311 | declined | not run |
| Score the exact shipped blend rule out-of-fold and allow fallback to the best validated single model | openml-training-333 | declined | not run |
pp = percentage points. Three deterministic repeats per executed case, not three independent tasks. Declined cases have no execution result. This mixed result does not establish robust transfer or CWM-agent efficiency gains.
The same frozen proposals were reviewed with public CSV statistics added. The new reviews were frozen before reading old execution results; no programs were rerun.
| Measured effect | Cases | Approved before | Approved with public profile |
|---|---|---|---|
| Improved | 4 | 4 | 4 |
| Unchanged | 14 | 14 | 7 |
| Worse | 4 | 4 | 4 |
| Not measured | 1 | 0 | 0 |
23 reviewed proposals; 1 excluded without a frozen review/code pair. One review per case, with reused baseline judgments: this is a context diagnostic, not a replicated improvement or a test of agent speed. Scope approval does not promise a score gain.
4 measured harmful changes, viewed in reverse: the edited program is weak and the original is strong. Original execution records and score directions stay unchanged.
4 lessons generated; 0 passed the unchanged source-quality audit. Audit rejections concerned unsupported causal claims and unresolved alternative explanations. No accepted library, cross-task test, or new code execution.
Same four training examples and outcome-blind auditor; only the generation instruction changes. Private outcome explanations must stay out of the served criterion. No quality gate is relaxed.
| Writing instruction | Lessons generated | Pass source audit |
|---|---|---|
| Original | 4 | 0 |
| Explicit separation of private evidence and criterion | 5 | 3 |
One generation per example. A source-audit pass is not evidence of transfer or an accepted library.
Blind static reviews on all four previously measured beneficial corrections, covering three other training tasks. No new program execution or development selection.
| Lesson | Favors better program | Favors worse program | No distinction / unknown | Supporting distinct tasks |
|---|---|---|---|---|
| Do not replace a default cutoff with an unsmoothed argmax over every observed probability value | 0 | 0 | 4 / 0 | 0 |
| Standardize each row when row-level scale dominates column-level differences | 0 | 0 | 4 / 0 | 0 |
| Require a noise margin or resampling check before adopting a tuned decision threshold | 0 | 1 | 3 / 0 | 0 |
24 / 24 program reviews usable. Pair counts are not independent tasks. This tests static discrimination on saved evidence, not the benefit of teaching the lesson to an agent. The 12-lesson library gate is unchanged.
Same 24 program/lesson inputs. The new arms use identical instructions, with either the target repeated or its counterpart shown as an unscored, non-citable comparison. Scores and outcome order remain hidden; lessons and citation checks are unchanged.
The saved-evidence audit found a curation defect: two threshold lessons require threshold tuning to apply, yet also accept avoiding it. The source audit missed this inconsistency. The diagnostic measures review consistency, not a repaired library.
| Review context | Usable reviews | Mixed scope | Favors worse program | Favors better program | No distinction / unknown |
|---|---|---|---|---|---|
| Single program (saved baseline) | 24 / 24 | 6 | 1 | 0 | 11 / 0 |
| Repeat target (control) | 21 / 24 | 4 | 1 | 0 | 8 / 3 |
| Other program (comparison) | 20 / 24 | 3 | 2 | 0 | 7 / 3 |
State: complete with unknown. 4 / 19 comparable target decisions differ between the two new arms; 5 targets lack a valid decision in at least one arm. Contrast columns count 12 lesson/correction combinations on just three tasks. Mixed scope means one program is applicable and the other is not; it is not automatically an error. Uncertain reviews remain separate.
Here, beneficial edits add the threshold tuning these lessons discourage. More reversals can expose a bad lesson rather than a bad reviewer. A favorable classification is not a success endpoint on this panel. The reused baseline also differs in framing and sampling; no agent, efficiency or medal gain is established. 48 new API requests maximum; no new code execution or cloud resources.
Rejected lessons do not end the whole study: the next planned source-expansion batch follows. The original quality checks and 12-lesson library gate stay unchanged. Expensive cross-task reviews run only after a source audit passes; source rejections stay in the denominator. Source-audited candidates in a diagnostic are not an accepted CWM library.
Each approved correction is compared with its unchanged original under the same execution cap, data, metric and seeds, with three repetitions. A case has six code runs: original and correction, each repeated three times. These repetitions are not independent tasks. Improved / unchanged / worse refers to balanced accuracy on fully paired cases; screened-out cases have no measured effect. Original interrupted results remain unknown; neither spent executions nor source programs are silently repeated. Final development selections are frozen before their execution scores are read; those scores cannot be used to rewrite or reselect lessons.
Deadline: 2026-10-05T17:33:00.144Z. At most four CPU Spot pods on MATS; 288 new program attempts, 300 seconds each. Owned pods are removed at completion or stop. This is small-tabular collection-pool development—not CWM-agent efficiency, medals, or a held-out benchmark result. No agent or benchmark arm is launched by this study.
29 measured collection-pool contrasts on 11 training tasks. Each fold excludes one whole task from generation. All conditions share source anchors, generation repetitions, models and validation. No benchmark or old development outcomes enter the study.
Single example uses one source pair. Multiple examples adds the other training evidence. Counterexample-aware uses the same additional evidence but explicitly reconciles improvements, regressions and unchanged results, and requires consistent scope for satisfying alternatives.
| Writing condition | Generation calls finished | Lessons generated | Source audit: pass / reject / unknown | Favors better / worse on excluded tasks | Unknown contrasts |
|---|---|---|---|---|---|
| Single example | 44 / 44 | 8 | 4 / 3 / 1 | 0 / 0 | 8 |
| Multiple examples | 44 / 44 | 10 | 5 / 2 / 3 | 0 / 0 | 3 |
| Counterexample-aware | 44 / 44 | 7 | 2 / 3 / 2 | 0 / 0 | 3 |
All review decisions are frozen. Unchanged-score cases and distinct-task support are reported separately from favorable ordering; missing responses remain unknown.
132 matched generation calls planned; at most 2580 model requests for the frozen design. Deadline: 2026-10-06T07:50:50.693106+00:00. Four concurrent calls; no automatic retries, new code runs or cloud resources. Scientific rejection advances the remaining groups. This is task-excluded TRAIN cross-validation, not a fresh benchmark result, an agent efficiency gain or an accepted library.
Compare each unchanged original with a Haiku task-only correction (no lesson) and a matched Haiku lesson-guided correction, selecting other TRAIN collection-pool programs marked violates before either correction or its outcomes. One lesson per correction: this tests lesson content, not the 12-rubric retrieval policy or a full agent episode.
complete; phase complete; updated 2026-10-06 01:59 UTC; deadline 2026-10-06T08:41:48.675293+00:00.
Request compatibility: 8 original requests returned HTTP 400; they remain failed requests, not negative lesson results. The native-schema compatibility check was provider_accepted in 9.42 seconds. The continuation reuses that response and runs only previously unstarted units. One intermediate schema-invalid response is also preserved, not repaired. The current wire schema is the exact native schema; explicit field instructions and native local validation remain. No deadline or overall request cap was increased.
Execution handoff: the first service could not find the cloud CLI tools. The corrected service environment passed a read-only MATS preflight. Execution reuses the frozen cases and approved programs; the earlier stop record is preserved. New model requests in this handoff: 0.
Model schedule scope_review: 28 / 28 finished; 0 active model calls. Requests: 286 reserved, 278 returned.
Curation: 9 grounded candidates from 9 planned generation requests; source audit 4 passed / 5 rejected or unknown (of 9); independent source consistency 4 / 4; 4 distinct lessons selected.
Target applicability: violates 59; satisfies 28; not applicable 76; insufficient evidence 0; unknown 18 (of 181 reviews). Selected: 19 cases on 10 tasks.
| Arm | Correction cases: approved / declined / rejected / unknown | Execution slots: attempted / planned | Slots: valid / invalid / unknown |
|---|---|---|---|
| Unchanged original | not a correction | 57 / 57 | 57 / 0 / 0 |
| Haiku, no lesson | 9 / 1 / 9 / 0 | 27 / 57 | 24 / 3 / 30 |
| Haiku, lesson-guided | 8 / 0 / 11 / 0 | 24 / 57 | 24 / 0 / 33 |
Comparison coverage: both corrections were approved in 4 / 19 cases across 4 tasks. Only paired-valid executions can establish their score difference; approval is not a score gain.
Local patch-validation failures before scope review: no lesson: 1; lesson-guided: 8. These are unusable proposed edits, not measured score losses.
Main contrast, lesson-guided minus no lesson: mean balanced-accuracy difference +0.0208; mean execution-seconds difference +0.7268. Paired-valid coverage 12 / 57; 45 pairs without both valid scores. Execution state: complete.
Same-case unchanged-program baseline
| Arm | Mean balanced accuracy | Mean program seconds |
|---|---|---|
| Unchanged original | 83.06% | 22.21 |
| No lesson | 80.82% | 21.95 |
| Lesson-guided | 82.90% | 22.68 |
Identical 4 cases across 4 tasks in all three rows; 12 dependent repetitions. A gain over no-lesson edits is not necessarily a gain over leaving the program unchanged.
Lesson-guided: 3 better / 1 worse cases than no lesson; 2 better / 2 worse than the unchanged original.
3 no-lesson timeout executions belong to 1 case(s) without an executed lesson-guided counterpart. They are not a paired validity advantage.
Measured writer defects (earlier study): 131 / 132 generations were schema-valid but yielded only 25 native candidates; 111 emitted lessons had empty applicability. Independent reviews: 82 / 99 usable, 17 unusable, including 13 truncated. Prospective repair uses explicit nonblank-field instructions, native local validation and a 6,000-token review cap; formatting repair is not evidence of better lessons.
Source-reviewed is not deployment-approved: native gates and the 12-lesson library gate remain unchanged. Effects condition on both arms valid; declined, rejected, missing and ambiguous outcomes are not zero scores. Three repeats preserve program seeds; repeats and cases sharing a task are dependent. TRAIN-only reused tasks, not a benchmark efficacy result, CWM-agent efficiency result or accepted library.
Envelope: 8 concurrent API requests / 4 GKE workers on MATS only; up to 24 cases × 3 arms × 3 repeats = 216 program attempts, 1,000 API requests and eight hours. Public JSON: plan · progress and results · writer diagnosis · request compatibility · correction coverage · matched final results.
Change under test: return one complete corrected program instead of substring replacements, in both the no-lesson and lesson-guided arms. Same frozen cases, lessons, models, output budgets, seeds and validation gates; the unchanged original remains the third arm. No selection of previous winners.
The previous study lost 8 / 19 guided proposals and 1 / 19 unguided proposals before scope review. This experiment first measures usable correction coverage, then score and runtime against both controls. A formatting fix is not evidence of better lessons.
Status: execution complete; interrupted outcomes retained; phase stopped; deadline 2026-10-06T15:14:05.190416+00:00.
Current model schedule: 29 / 29 finished; 0 active. Total requests: 69 reserved, 69 returned.
Execution-only recovery after infrastructure stops. 1 previously unstarted executions; the earlier 114 valid executions and 4 interrupted/unknown outcomes are retained. 0 new model requests. No attempted program is retried; the original deadline and scientific gates are unchanged.
1 earlier invalid execution is also retained, not retried.
Finish the sole unstarted slot after the disk-reserve stop; retain every previous outcome.
Saved-output audit: three file-reference rejections came from changed module docstrings containing existing paths, not new file access. Three other outputs do not parse as Python; one uses blocked dynamic class construction. Original rejections stay in place. These checks do not establish whether the rejected edits would improve predictions.
| Arm | Cases: approved / declined / rejected / unknown | Executions: attempted / planned | Valid / invalid / no observed outcome |
|---|---|---|---|
| Unchanged original | not a correction | 57 / 57 | 56 / 0 / 1 |
| No lesson | 8 / 0 / 11 / 0 | 24 / 57 | 21 / 1 / 35 |
| Lesson-guided | 13 / 0 / 6 / 0 | 39 / 57 | 38 / 0 / 19 |
No observed outcome includes unapproved and unstarted slots, as well as interrupted executions.
Local construction/validation failures, No lesson: introduced_data_or_file_reference: 1; python_parse_error: 1.
Local construction/validation failures, Lesson-guided: introduced_data_or_file_reference: 2; introduced_or_changed_external_access: 1; python_parse_error: 2.
Same-case comparison
| Arm | Mean balanced accuracy | Mean program seconds |
|---|---|---|
| Unchanged original | 77.16% | 6.48 |
| No lesson | 73.82% | 6.73 |
| Lesson-guided | 77.88% | 10.30 |
Identical 5 paired-valid cases across 5 tasks; 15 dependent repetitions. Means average each case's repetitions first; repeated executions are not new tasks.
Partial-evidence comparison: interrupted outcomes remain unknown and are excluded only from paired-valid means, never from the all-case accounting above.
Stop reason: all_admitted_slots_exhausted_partial_evidence.
All selected cases stay in the denominators. Rejected, declined and missing results are not zero scores. TRAIN-only diagnosis, not a medal-rate or full-agent efficiency result. No library deployed and no acceptance gate relaxed. Limit: 76 model requests, 8 concurrent; 4 MATS CPU workers, 300 seconds per execution, 3 repetitions and an eight-hour deadline. Public JSON: plan · progress and results · saved-output diagnosis.
| curation method | rubrics | runs | valid | medal | above median | valid submissions per run (of 22) | library vs execution | shown vs execution | library vs pool |
|---|---|---|---|---|---|---|---|---|---|
| References (no CWM) | |||||||||
| Ceiling | – | 4 | 75% (16.5) | 8% | 18% | 18 / 16 / 16 / 16 | – | – | – |
| Floor | – | 4 | 44% (9.8) | 5% | 9% | 7 / 15 / 9 / 8 | – | – | – |
| Hybrid CWM · Retrieval: Similarity (the 12 most similar rubrics from the whole library). Rows marked reference are the oracle libraries, cut to a known failure mix | |||||||||
| Wrong-answer | 1,000 | 4 | 58% (12.8) | 8% | 12% | 14 / 13 / 13 / 11 | 89 | 85 | 78 |
| Unmatched | 4,294 | 4 | 57% (12.5) | 8% | 9% | 11 / 13 / 13 / 13 | 73 | 86 | 62 |
| Filtered | 220 | 4 | 67% (14.8) | 5% | 10% | 12 / 16 / 15 / 16 | 45 | 46 | 4 |
| Quota | 952 | 4 | 73% (16.0) | 7% | 11% | 15 / 16 / 15 / 18 | 48 | 60 | 4 |
| Oracle, pool-matched (reference) | 187 | 4 | 83% (18.2) | 8% | 11% | 20 / 17 / 19 / 17 | 45 | 54 | 27 |
| Oracle, MLE-bench-matched (reference) | 186 | 4 | 68% (15.0) | 5% | 7% | 14 / 17 / 12 / 17 | 13 | 44 | 50 |
| Hybrid CWM · Retrieval: Category split (12 slots divided across failure categories, then most similar within each) | |||||||||
| Filtered | 220 | 4 | 75% (16.5) | 7% | 12% | 18 / 18 / 16 / 14 | 47 | 49 | 2 |
| Quota | 952 | 4 | 85% (18.8) | 5% | 9% | 20 / 16 / 20 / 19 | 48 | 54 | 4 |
Rates are over all finished runs of 22 competitions each (mean valid submissions in brackets). Runs of the same setup vary a lot (the floor: 7, 15 and 9 valid), so gaps of a few points are noise. Mismatch columns: total variation distance on the 22 failure categories with crashes split by type, 0 = identical mix, 100 = no overlap. Execution = the failures MLE-bench Lite shows when every command runs (three ceiling runs: failed commands plus the cause of each wrong result); pool = the collection pool's failures, which the libraries were matched to. Hover a header for its definition.
| row | recorded window / competition | ceiling / row time | real execution during episode |
|---|---|---|---|
| Ceiling | 55.2 min | 1.0× | 39.9 min |
| Floor | 21.6 min | 2.6× | 0.0 min |
| Wrong-answer · Similarity | 21.9 min | 2.5× | 6.1 min |
| Filtered · Similarity | 24.5 min | 2.3× | 6.7 min |
| Filtered · Category split | 21.5 min | 2.6× | 7.7 min |
| Quota · Similarity | 25.2 min | 2.2× | 5.8 min |
| Quota · Category split | 26.2 min | 2.1× | 10.0 min |
Historical records lack exact episode boundaries. This window includes logged setup waits and the final collection run, but not unlogged provisioning, grading or earlier retry attempts; it is not pure agent time. Means include valid and invalid episodes. Real execution excludes installs and read-only harness probes. The hybrid changes execution policy as well as feedback; this ratio does not isolate the rubric contribution.
| row | CWM predicts | library rubrics | real code runs during the episode | feedback: CWM only / real output | real run min / competition | runs | valid | medal | above median | valid per run (of 22) |
|---|---|---|---|---|---|---|---|---|---|---|
| Floor | no | no | none | 0 / 0.0 | 0.0 | 4 | 44% (9.8) | 5% | 9% | 7 / 15 / 9 / 8 |
| Real run at submit only | no | no | at submit | 0 / 0.7 | 1.2 | 4 | 55% (12.0) | 7% | 9% | 10 / 10 / 15 / 13 |
| Checks on every change | no | no | every code change (final.py, 3 min cap) and at submit | 0 / 12.9 | 18.7 | 4 | 69% (15.2) | 8% | 14% | 15 / 14 / 15 / 17 |
| Pure CWM, no rubrics | yes (task text only) | no | none | 5.4 / 0.0 | 0.0 | 4 | 38% (8.2) | 6% | 9% | 8 / 8 / 7 / 10 |
| Pure CWM | yes | Quota, Category split | none | 8.4 / 0.0 | 0.0 | 4 | 53% (11.8) | 7% | 11% | 11 / 11 / 13 / 12 |
| Hybrid CWM, no rubrics | yes (task text only) | no | on clean predictions and at submit | 4.0 / 4.3 | 6.2 | 4 | 75% (16.5) | 6% | 11% | 15 / 14 / 19 / 18 |
| Hybrid CWM | yes | Quota, Category split | on clean predictions and at submit | 5.7 / 6.9 | 10.0 | 4 | 85% (18.8) | 5% | 9% | 20 / 16 / 20 / 19 |
| Ceiling | no | no | every code command, 600s per-command limit | 0 / every run | 39.9 | 4 | 75% (16.5) | 8% | 18% | 18 / 16 / 16 / 16 |
Same 22 competitions, same agent. The real run at submit is capped at 3 min and only holds a submission whose final.py crashes or that the CWM still flags. 'No rubrics': the CWM judges the change against the task statement alone (no library). The earlier scheduled-check sweep was stopped after GKE transport errors entered agent feedback; its partial grades are not included here or used for causal inference. It also differed in CWM submit vetoes. The fixed-policy controls below remove that veto and add a task-only CWM control. Feedback counts cover messages shown to the agent, not silent accepted-submit checks. Real-run time excludes environment/data inspection, installs and final collection. 'Checks on every change' runs the deliverable on every routed code request, without a CWM call.
Same execution policy, fixed-schedule checks (defined in the settings table): the agent's 2nd, 4th, 6th… code-running commands each trigger a real run of final.py, stopped after at most 180 seconds; the other code commands are not executed. This counts commands; it is not a timer. At submit final.py runs again; only a real crash holds the submission (at most twice), and the CWM has no submit veto. The two CWM rows also show the CWM's prediction on every code command. Same agent and limits; Quota uses the frozen 952-rubric library, with no new mining. Four full 22-competition runs per row.
| feedback | graded | valid | medal | above median | real runs | real min | episode min | episode + final run min | lost finals |
|---|---|---|---|---|---|---|---|---|---|
| Quota rubrics + CWM | 88/88 | 78.4% | 8.0% | 13.6% | 6.0 | 7.7 | 14.7 | 21.6 | 0 |
| Task-only CWM | 88/88 | 64.8% | 6.8% | 13.6% | 6.0 | 8.6 | 15.1 | 23.8 | 6 |
| Checks only, no CWM | 88/88 | 55.7% | 2.3% | 11.4% | 7.6 | 10.9 | 17.4 | 26.0 | 5 |
Collection-loss caveat: all episodes remain in the denominator. Lost finals are not proven code failures; the observed CIs are not clean causal estimates. Sensitivity changes only missing final validity, not runtime. Nine agent-loop OOM terminations per row remain real failures. Times cover the last saved attempt, including failed outcomes and CWM latency; final collection is included in the last time column, but earlier retries, provisioning and grading are not. CIs resample competitions, keeping four repeats together. Same check rule does not force identical realized check counts or duration.
What this tests: Quota versus task-only isolates adding the curated library under this policy; Quota versus checks-only tests the whole feedback channel. Higher validity at longer runtimes is a tradeoff, not an efficiency win. These controls do not by themselves prove adaptive triggering or distribution matching is the cause.
Reading: the package has a validity advantage over checks even if their five lost finals all succeeded. The extra benefit of curated rubrics over task-only CWM is less certain after collection-loss sensitivity. A rubric-specific speedup is not established; medals are 7 versus 6, and above-median outcomes are 12 versus 12.
Quality lessons exist, but this category is not retrieved. S9P01 is the existing category “Modelling choice leaves score on the table”.
| stage | performance-gap category / total | share |
|---|---|---|
| Pool failure-event reference | 138 / 4,828 | 2.9% |
| Curated library | 29 / 952 | 3.0% |
| Library support used for slot allocation | 29 / 1,747 | 1.7% |
| Allocated retrieval slots | 0 / 12 | 0.0% |
| Logged library retrieval exposures | 0 / 11,088 | 0.0% |
Its support gives 0.20 of 12 slots before integer allocation, then 0 afterwards. The fixed split spends ten slots on crashes, timeouts and missing packages, one on output format and one on missing-value checks. Across 924 saved verdicts, S9P01 was never retrieved, marked applicable or flagged. Baseline-validation, metric-selection and hyperparameter-validation categories also had zero retrieved exposures.
Logged verdicts are not all grade-wrapper attempts or cached feedback deliveries. Other categories and the task criterion can still contain modelling advice. This establishes a coverage gap, not that filling it will win medals; training limits, CWM accuracy and agent use of advice remain untested.
Previous validation collection — probe: 64/64 new OpenML attempts recorded; 64 valid. Eight new small classification tasks, eight real interpreter attempts each. These tasks are reserved for validation, not rubric mining.
The frozen CWM comparison used current retrieval, two protected quality slots, and the task criterion alone. The protected-slot screen did not clear its complete-coverage gate, so it did not launch a Lite arm. The new curation experiment above uses separate training tasks, not these validation examples.
The existing collection had no unused task-disjoint examples. OpenML tests small-tabular quality discrimination, not medal gains on large competitions. Its datasets are public. Final collection now uses a durable exit record; a MATS smoke passed and its pod was deleted. New retries preserve all attempts rather than overwriting their costs.
| row | rubrics | distance to the pool's mix | runs | valid | medal | above median | valid per run (of 22) |
|---|---|---|---|---|---|---|---|
| Ceiling | 4 | 75% (16.5) | 8% | 18% | 18 / 16 / 16 / 16 | ||
| Floor | 4 | 44% (9.8) | 5% | 9% | 7 / 15 / 9 / 8 | ||
| Quota · Category split, hybrid CWM (952) | 4 | 85% (18.8) | 5% | 9% | 20 / 16 / 20 / 19 | ||
| Pure CWM, Quota planned 10 | 8 | 24 | 4 | 35% (7.8) | 2% | 8% | 6 / 8 / 7 / 10 |
| Pure CWM, Quota planned 100 | 99 | 4 | 4 | 56% (12.2) | 8% | 12% | 12 / 12 / 12 / 13 |
| Pure CWM, Quota planned 500 | 500 | 1 | 4 | 38% (8.2) | 7% | 8% | 6 / 8 / 11 / 8 |
| Pure CWM, Quota planned 1,000 | 952 | 4 | 4 | 53% (11.8) | 7% | 11% | 11 / 11 / 13 / 12 |
| Pure CWM, Quota planned 2,000 | 1,802 | 8 | 4 | 55% (12.0) | 7% | 10% | 11 / 11 / 12 / 14 |
Pure CWM: no code runs for real during the episode (no run on a clean prediction, none at submit); every library is curated from scratch by the Quota miner at that size (planned size; a category whose new failures were already covered stops early, so large sizes can come out smaller), served by Category split. The 1,000 row is the method's own library (952).
| row | agent budget | runs | valid | medal | above median | valid per run (of 22) |
|---|---|---|---|---|---|---|
| Ceiling | 2 h, 250 steps | 4 | 75% (16.5) | 8% | 18% | 18 / 16 / 16 / 16 |
| Ceiling | 4 h, 500 steps | 2 | 75% (16.5) | 7% | 16% | 17 / 16 |
| Quota · Category split, hybrid CWM | 2 h, 250 steps | 4 | 85% (18.8) | 5% | 9% | 20 / 16 / 20 / 19 |
| Quota · Category split, hybrid CWM | 4 h, 500 steps | 2 | 82% (18.0) | 5% | 11% | 17 / 19 |
Doubled wall clock, step and cost limits; the per-command time limit (10 min) is unchanged.
| all competitions | the 22 of Lite | the 53 outside Lite | |||||||
|---|---|---|---|---|---|---|---|---|---|
| row | valid | medal | above median | valid | medal | above median | valid | medal | above median |
| Ceiling | 75% (69) | 3% | 10% | 82% (22) | 0% | 14% | 72% (47) | 4% | 9% |
| Floor | 49% (69) | 1% | 4% | 41% (22) | 5% | 5% | 53% (47) | 0% | 4% |
| Quota · Category split, hybrid CWM (952 rubrics) | 84% (69) | 3% | 6% | 82% (22) | 5% | 9% | 85% (47) | 2% | 4% |
One run per row of the full 75-competition MLE-bench, same harness as Lite (CPU pods, 2 h agent budget), counted on the competitions graded in all three rows (in brackets). Left out: three whose mle-bench graders fail on pandas 3 (tgs-salt, tensorflow2-question-answering, vinbigdata) and three whose ceiling pods were evicted in every attempt because the agent unpacked the data past the pod's 10 GB disk (freesound-audio-tagging-2019, inaturalist-2019, iwildcam-2019; the floor and the CWM row had valid submissions on the first two, so leaving them out favours the ceiling). The 53 competitions outside Lite were never used for error analysis or any choice: the first clean test of the method.
| row | rubrics | run 1 | run 2 | vs floor, paired (95% CI) | papers better / worse |
|---|---|---|---|---|---|
| Ceiling | – | 0.104 | 1/20 graded; grader failed on 19 | -0.000 [-0.035, +0.032] | 7 / 12 |
| Floor | – | 0.105 | 0/20 graded; grader failed on 20 | – | – |
| Crash-rubric · Similarity (MLE-bench pool) | 963 | 0.147 | +0.043 [+0.010, +0.078] | 14 / 5 | |
| Quota · Category split (PaperBench's own pool) | 628 | 0.115 | +0.011 [-0.009, +0.030] | 12 / 8 | |
| Quota · Category split (MLE-bench pool, transfer) | 952 | 0/20 graded; grader failed on 5, 15 not run | – | – |
Mean score over the 20 papers of the all split (90 min agent, 1 h reproduction), every arm graded by Claude Sonnet 5.5 (earlier arms regraded from their stored submissions). A difference is real only when its interval excludes 0. The Crash-rubric library was picked for PaperBench because its failure mix was closest to PaperBench's own (distance 35 vs 56 for PaperBench's pool library): that row has looked at the benchmark.
No mismatch chart here: PaperBench's collection pool and its benchmark were labelled in different category sets.
Both benchmarks have been used to find what goes wrong, so these are development-set numbers.