Rubric-grounded code world models — evaluation roadmap

MLE-bench, KernelBench and SWE-bench are fully held out; rubrics are mined only from external datasets. What is measured, what is built, what is next.

1Standard evaluation practice for MLE agentsbackground

MLE-bench is the reference protocol. An agent receives a Kaggle competition's public data and description and must produce a submission.csv; that file is scored by the competition's own metric against a held-out private split, and the score is placed on the competition's real historical leaderboard. The headline number is medal rate — the share of competitions clearing the actual bronze, silver or gold thresholds, which are set as leaderboard percentiles so a medal means comparable achievement whether a competition drew 200 entrants or 5,000. Above-median rate is the usual secondary figure. Attempts run under a fixed wall-clock and hardware budget with no internet access, and because agent variance is large, results are reported over multiple seeds as pass@1 or pass@k: the strongest published setup, o1-preview with AIDE scaffolding, earns bronze or better in 16.9% of competitions at pass@1 and 34.1% at pass@8. Grading is always done by the released grader rather than a re-implementation, and a 22-competition "Lite" split exists for cheaper runs. Contamination is an acknowledged weakness, since the competitions predate current model cutoffs.

2Benchmarks and mining datasets, by disciplineall datasets

Three disciplines, each with one benchmark held out for evaluation and the rest used only for mining.

Every mining source passes an explicit leakage filter, and the counts below are post-filter. The filter is not a formality — it caught real overlap in all three domains:

The filter runs at three points: when a pool is enumerated, again at collection time via the split guard, and once more at database build, where every rubric is traced back to its source task and the build fails if any traces to a held-out task.

disciplinebenchmarktasksroletask source
ML engineering
& data science
MLE-bench75held outhuman · Kaggle
DA-Code500mininghuman
InfiAgent-DABench257mininghuman CSVs · synthetic questions
DSBench66mininghuman · Kaggle, re-split
MLGym19mininghuman
MLAgentBench14mininghuman
RE-Bench7mininghuman · expert-authored
GPU kernels KernelBench270held outhuman
TritonBench55mininghuman
FastKernels47mininghuman · production archs
GPU MODE reference-kernels42mininghuman · competition
BackendBench26mininghuman · PyTorch aten ops
robust-kbench11mininghuman
Software engineering SWE-bench12 reposheld outhuman · real issues
SWE-smith131 reposminingsynthetic · LM-injected bugs
SWE-Gym11 reposmininghuman · real issues/PRs

1,176 mining units in total. Everything is human-authored except SWE-smith, the one large synthetic block, which matters because LM-injected bugs need not resemble real failure modes. Held-out sizes are the published figures; mining counts are our own enumeration of each repository at a pinned commit and can differ from a paper's headline number, since a paper may count individual operators or questions where the repo ships grouped suites. SWE rows are repositories, each yielding many task instances at collection time. Full inventory as an image →

3End-to-end accuracydata: MLE-bench · in-distribution

These rubrics were mined from MLE-bench competitions, so this run is in-distribution: the mining and eval competitions are disjoint, but both come from MLE-bench. It is a check that the end-to-end loop runs and is measurable, not a held-out result.

Setup

Each trial takes one held-out MLE-bench program that has already been executed and graded, hands it to a coding agent, and gives that agent 30 minutes of wall clock to produce better code. Five arms differ only in what feedback the agent gets each turn: revise gets none (the floor), interpreter executes the code and reads the real traceback, world model retrieves rubrics and has a reward model judge which are violated without ever executing, WM + patch applies the rubric's suggested edit directly, and race runs the world model against a real execution and uses whichever answers first. The clock is the only budget, so a fast feedback channel buys more turns — the world model gets roughly 50 turns where the interpreter gets 3–15. At the deadline each arm submits its best version chosen by its own signal, never by the true score, and that submission is executed once and graded by the official MLE-bench grader outside the budget.

Wall-clock e2e results by scaling stage

Reading the figure

Each group on the x-axis is one rubric database, labelled by its mining/eval competition split and its error and performance rubric counts; the groups run from a 151-rubric database up to 708. Within a group the five bars are the five arms. The y-axis is improvement rate: the percent of trials where the submitted code scored better than the seed program it started from. A dashed rule separates an earlier old-library run kept for reference, whose competition set differs from the staged ones.

Last Saturday's end-to-end run, using the rubrics presented that day, finished two days ago. We are lagging slightly behind interpreter-only, controlling for wall clock so every arm gets 30 minutes. The higher numbers in the earlier stages are mainly an artifact of having fewer eval tasks there, so the percentage shifts easily with run-to-run noise.

The x-axis groups runs by rubric-database size, each group showing the five arms. The y-axis is improvement rate: the percent of attempts where the final code scored better than the seed starting program, an existing MLE-bench code trace shared by all runs. MLE-bench officially reports medal rate — how many competitions the agent's submission clears the Kaggle leaderboard's bronze, silver or gold thresholds — but our 30-minute runs do not score enough medals on the smaller split to measure improvement, so improvement rate is the better gauge of whether the end-to-end loop is working. For the full run, once MLE-bench is completely held out, we can use medal rate and the official metrics.

4Rubric precision / recall, updated generation pipelinedata: MLE-bench · in-distribution

These rubrics are also mined from MLE-bench competitions, so they are in-distribution in the same sense as the section above.

After these end-to-end results, William and I looked through the traces from the run and found coverage issues coming from the prompts and the generation pipeline. William improved the pipeline by asking the rubric-generation model to focus on coverage and by adding filtering steps, which gave much higher P/R. Below is the new rubric database's P/R. We are currently running an end-to-end experiment on these updated, better-coverage rubrics. Even with the better generation, the performance rubrics appear saturated.

Setup

This measures the rubrics themselves rather than an agent. Each rubric is shown to a judge together with a held-out program and the static facts available before that program ran — the data files, their columns, the installed packages — but never the program's output. The judge decides whether the rubric fires. Truth comes from what actually happened when the program was executed: for the error family, whether the interpreter raised the exception class the rubric predicted; for the performance family, whether annotators credited that criterion for a measured score gap between two programs. Timeouts and environment kills are excluded, since no static rubric can predict them.

Precision and recall of the updated rubric database

Reading the figure

Precision is the share of firings that were correct — how often a rubric that flagged a program was right about it. Recall is the share of real failures that some rubric caught. Both are reported separately for the error and performance families, since they answer different questions and are judged against different truth.

5Curation speeddata: MLE-bench + external

regimeerror / pod-hrperf / pod-hrcombined
MLE-bench (3600 s cap, ≤40 programs) 1.03.54.5
External (DSBench + kernels, 900 s cap, ≤8 programs) 20.042.963.0

Where the time goes

componentsecondsshare of wall
GPU execution of candidate programs140,67494%
LLM generation + curation + overhead8,8846%

Curation is not the bottleneck — execution is. Median 267 s of GPU time per program and 302 s per rubric, but spread over two orders of magnitude:

competitionGPU s / rubricrubrics
random-acts-of-pizza2622
statoil-iceberg5430
stanford-covid-vaccine30236
dog-breed-identification1,11622
nyc-taxi-fare1,75518
vesuvius-ink-detection3,8118

Four of eleven runs (AI4Code, osic, uw-madison, vesuvius) spent 34,787 GPU-seconds for zero performance rubrics: no program was graded twice, so no better/worse pair exists. 27 runs were launched, 11 banked a summary, none reached its rubric target.

Scaling levers, by measured payoff

levereffectbasis
Re-mine the 3,355 already-executed states on disk~20–50× costs only the 6% slice; 1,431 error states, 802 graded, corpus not exhausted
Mine external tasks instead of MLE-bench14× 63.0 vs 4.5 rubrics/pod-hr, measured
Cap mining execution (~300 s + subsampling)~10× crashes surface at a median 11 s; needs a fidelity check for perf pairs
Abandon a competition with no graded pair after N minreclaims 23% 34,787 of 149,558 s produced no perf rubric
Cross-program perf pairs instead of hill-climbingremoves 1 run/turn 802 graded states already exist; weaker evidence, needs measuring
More podslinearone competition per pod

61,000 rubrics uniformly across the mining pooldata: external mining poolrunning

Curate 1,000 rubrics distributed uniformly over all thirteen mining datasets and evaluate, producing a scaling-law figure with rubric count on the x-axis and end-to-end accuracy on the y-axis, plus companion panels for error and performance precision and recall. Two things this run has to do that the current experiment cannot: hold the eval set fixed across rubric counts, since the present design varies library size and task difficulty together and so has no real scaling axis; and capture medal rate and above-median, which the harness computes but which were missing from the stage campaign because the medal fields postdate that code snapshot and the per-attempt rows were dropped by an artifact-glob bug. Both are fixed, so a rerun yields the metric comparable to published MLE-bench numbers.

7Collect externally, evaluate held outdata: external → held-out benchmarksrunning

Collection now runs end to end on the external domains, verified on DSBench (five and four performance rubrics on two tasks) and on kernels (six), with error rubrics in both. The next step is to collect rubrics from the other sources in the table above — DA-Code, InfiAgent-DABench, DSBench, MLGym, MLAgentBench, RE-Bench, TritonBench, FastKernels, GPU MODE reference-kernels, BackendBench, SWE-Gym and SWE-smith — and then run end-to-end on those, building rubric_db_mle, rubric_db_kernel and rubric_db_swe; audit each database with build_domain_db.py, which traces every rubric to its source task and fails the build on a held-out origin; then evaluate on the fully held-out benchmarks, giving leaderboard-honest numbers for the first time. Two gaps to settle first are measurement decisions rather than plumbing: kernel and SWE failures surface inside the grader rather than the interpreter, so error-rubric yield in those domains is currently limited to script-level defects, and SWE error truth needs extending from exception class to failing-test identity.