← back to the main dashboard · curation cost per rubric

MLE-bench Lite

MLE-bench Lite, all 21 competitions held out · Haiku 4.5 agent · official mlebench grader · 2-hour budget per competition · rubrics collected from external data-science task pools only (270 pool rollouts, 69 failures, 69 lessons: the library is 10 / all = 69, and the "100" arms are replicates of "all")

Superseded. This page reports the 9/07 run-and-hide design: execution commands were still run in the pod and only their output was hidden, so wall clock here includes real training/test execution, the side-channel scans after every execution, and (for some arms) API throttling. The canonical never-execute results are on MLE-bench Lite — canonical hybrid CWM.

1Experiment setting

Each task is a Kaggle competition: the agent gets the competition description and data and must leave a valid submission file; the official grader scores it against the real leaderboard. Medal = the submission clears the bronze threshold; the table also reports the above-median rate and valid submissions.

A Haiku 4.5 agent works in the repository by issuing bash commands. The settings differ only in what answers those commands:

In every world-model setting an execution command is still run in the pod (its side effects persist) and only the observation the agent sees is replaced by the world model's answer; reads, searches and edits are always real. Same settings as the 9/07 SWE-bench page.

Interpreter baseline — every command runs; the agent sees real output.
agentHaiku 4.5bash commandinterpreterruns everything, 60 s capreal output + return code
No code execution (floor) — reads and edits are real; runs return nothing.
agentHaiku 4.5routerrule: does it run code?reads / edits / grepinterpreterreal output + rcreal outputcode executionnothing backoutput withheld · rc hiddenthe run still happens; files it writes are quarantined
All-CWM (original) — the CWM answers everything, using the same retrieved rubrics; the agent cannot read code.
agentHaiku 4.5every command, reads includedCWMgrades the diff vs the retrieved rubricsrubric verdict onlyrubric library10 / 100 / 1000 mined rubricstop-12 nearest to the taskthe agent never sees file contentsor any command's output
Hybrid CWM (this work) — reads go to the interpreter, runs go to the CWM with retrieved rubrics.
agentHaiku 4.5routerrule: does it run code?reads / edits / grepinterpreterreal output + rcreal outputcode executionCWMgrades the diff vs the retrieved rubricsrubric library1000 mined rubricstop-12 nearest totask + diff + commandverdict + real rcthe run executes for side effects only; its output is never shown

The router is a fixed rule, not a model. The command string is split on every bash separator (\n ; && || | &), heredoc bodies are dropped and wrapper prefixes (sudo, timeout N, VAR=x) skipped; a command is code execution if any segment's first word is an interpreter or test runner (python, pytest, tox, make, manage.py, runtests.py, bash, sh, …), a script or repo executable (./run.sh, /testbed/bin/x), or contains a substitution / xargs / find -exec that runs one. Everything else — cat, ls, grep, sed -i, git, heredoc writes — is a read or edit and goes to the interpreter. A mixed command such as cat > t.py <<EOF … EOF && python t.py counts as execution.

The CWM (Haiku 4.5 or Opus 5) reads the task, the agent's current diff and the command, grades the diff against the 12 rubrics retrieved from the mined library (nearest to the task, the files the diff touches and the command; in the hybrid arms a standing task criterion — “does the change implement what the issue asks?” — is added to the list) and returns the per-rubric verdict. Files written by a run are quarantined so the agent cannot read the output back; every trajectory is audited for leakage. No setting ever sees the hidden grading tests.

2Results

any medal, % of 22

0%12%24%36%48%60%14%3/220%0/220%0/220%0/220%0/210%0/220%0/210%0/220%0/220%0/21pending5%1/225%1/210%0/220%0/22interpreterrealexecutionno code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

above the leaderboard median, % of 22

0%12%24%36%48%60%32%7/229%2/220%0/220%0/220%0/215%1/225%1/215%1/2214%3/225%1/21pending5%1/2214%3/215%1/220%0/22interpreterrealexecutionno code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

MLE-bench's headline metric is the medal rate (top chart); "above the leaderboard median" is the secondary, easier bar. Valid-submission counts are in the table.

above the leaderboard median, % of 22

0%12%24%36%48%60%32%7/229%2/220%0/220%0/220%0/215%1/225%1/215%1/2214%3/225%1/21pending5%1/2214%3/215%1/220%0/22interpreterrealexecutionno code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc
armcountratealso
interpreter · real · execution3/2213.6%above median 7 · valid 18
no code exec. · bash · only0/220.0%above median 2 · valid 16
all-CWM · Opus · 10 rubrics0/220.0%above median 0 · valid 0
all-CWM · Opus · 100 (69-rubric library)0/220.0%above median 0 · valid 0
all-CWM · Opus · all (69-rubric library)0/210.0%above median 0 · valid 0
hybrid CWM · Haiku · task only0/220.0%above median 1 · valid 16
hybrid CWM · Haiku · all rubrics0/210.0%above median 1 · valid 12
hybrid CWM · Opus · task only0/220.0%above median 1 · valid 15
hybrid CWM · Opus · 10 rubrics0/220.0%above median 3 · valid 19
hybrid CWM · Opus · 100 rubrics0/210.0%above median 1 · valid 14
hybrid CWM · Opus · 1000 rubricspending
hybrid CWM · Opus · all rubrics (69-rubric library)1/224.5%above median 1 · valid 18
hybrid CWM · Opus · 12 random (69-rubric library)1/214.8%above median 3 · valid 13
simulator · Haiku · pure0/220.0%above median 1 · valid 9
simulator · Haiku · + real rc0/220.0%above median 0 · valid 6
hybrid CWM · Opus · 6 generated1/224.5%above median 3 · valid 16

Efficiency: median wall-clock per instance

0.0 min26.0 min52.0 min78.0 min104.0 min130.0 min52.5 min65.6 min9.2 min14.9 min14.0 min47.7 min47.5 min30.3 min16.4 min15.9 minpending31.5 min23.8 min75.8 min35.5 mininterpreterrealexecutionno code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

Wall-clock per instance from the first command to the last (minutes), from the per-command timing records of each pod.

Efficiency: median agent LLM calls per instance

06012018024030046913443357763656358pending605812794interpreterrealexecutionno code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

Efficiency: median seconds per LLM call (wall clock / calls, per instance)

0.0 s12.0 s24.0 s36.0 s48.0 s60.0 s51.4 s24.2 s15.6 s39.5 s17.7 s30.4 s37.8 s19.5 s15.3 s16.4 spending22.2 s18.3 s28.0 s18.2 sinterpreterrealexecutionno code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

Wall clock is calls × seconds per call. This chart isolates the second factor: the time the agent waits per step, which contains the model call and, in the world-model arms, the judge verdict; in the interpreter it contains the real execution. An arm that is slower here costs more per step; an arm that is slower only in wall clock took more steps.

Efficiency: median cost per instance (agent + world model, $; judge cost = graded observations x that arm's $/grade)

$0.00$1.20$2.40$3.60$4.80$6.00$0.38$0.47$8.00$11.61$9.45$0.77$1.14$5.97$7.55$7.94pending$9.16$8.87$0.64$0.66interpreterrealexecutionno code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

Mock: median wall clock per competition under the canonical architecture (code execution never runs)

0.0 min16.0 min32.0 min48.0 min64.0 min80.0 min54.2 min68.4 min35.7 min9.6 mincanonical mock8.0 minmock9.9 minmock11.9 minmockinterpreteras runno code exec. flooras runhybrid Opus allrun-and-hideas runhybrid Opus allexecution removedinterpreterexec cmds → judge10 sinterpreterexec cmds → judge15 sinterpreterexec cmds → judge20 s

Computed from the recorded per-command timings of the runs above. "Execution removed" subtracts the time of every command the router classified as code execution from the hybrid arm's wall clock (its judge calls are already inside that wall clock). The interpreter rows replace each of its execution commands by one judge call of 10 / 15 / 20 s. Totals over the set: interpreter 23.6 h, hybrid as run 18.0 h, hybrid with execution removed 3.9 h (83% below the interpreter). These are latency mocks: without execution nothing writes the submission file, and agents behave differently when they never see a training result.

What the agent does: command mix per setting

Every agent command classified by the routing rule (inner ring: interpreter vs code execution) and by type (outer ring). "Write file" = heredoc / echo redirects; "edit / file ops" = sed -i, mv, cp…; "run tests" = pytest and friends; "python script" = python x.py; "inline python" = python -c / heredoc.

interpreter — every command real

pending

no code execution — runs return nothing

pending

hybrid CWM · Opus · all — runs answered by the Opus CWM

pending

hybrid CWM · Haiku · all — runs answered by the Haiku CWM

pending

3Example traces

Trace pair pending: needs the same instance in mle-cwmall and mle-noexec.

4Curation cost per rubric

On-policy mining from the DABench + DA-Code data-science pool (run_da_pool.py): the same Haiku agent runs on pool tasks with real execution, every failed rollout is distilled into one lesson by Opus, lessons are deduplicated and cut into the 10 / 100 / all libraries. No CWM grading during mining and no precision filter, so the per-rubric cost is the "cheap" recipe of the SWE-bench page. Prices: Opus 5 $5 / $25 per M tokens, Haiku 4.5 $1 / $5, GKE Spot ≈ $0.045 per node-hour.

itemvalue
pool rollouts (Haiku interpreter agent, real execution)pending
median agent cost per rolloutpending
failures distilled (Opus, one call per failure, ~$0.09 each)pending
lessons after dedup at 0.90 = library sizepending
cost per kept rubric (agent + distillation + pods)pending