← back to the main dashboard · curation cost per rubric

MLE-bench Lite — harness 2 (9/14): interpreter, CWM-only, dual channel

MLE-bench Lite, all 22 competitions held out · Haiku 4.5 agent · official mlebench grader · 2-hour budget per competition · harness 2 (pulled 2026-09-14): agent told its channel, installs never routed, judge latency recorded; every arm writes final.py which the harness runs once, uncapped, at collection. CWM library = on-policy all + task criterion + preamble. Dual arms run each routed command for real under a 30 / 60 / 120 s budget AND grade the change set.

Superseded. This page reports the 9/07 run-and-hide design: execution commands were still run in the pod and only their output was hidden, so wall clock here includes real training/test execution, the side-channel scans after every execution, and (for some arms) API throttling. The canonical never-execute results are on MLE-bench Lite — harness 2 (9/14): interpreter, CWM-only, dual channel — canonical hybrid CWM.

1Experiment setting

Each task is a Kaggle competition: the agent gets the competition description and data and must leave a valid submission file; the official grader scores it against the real leaderboard. Medal = the submission clears the bronze threshold; the efficiency page uses 'beats the no-execution floor' as the solved bar.

A Haiku 4.5 agent works in the repository by issuing bash commands. The settings differ only in what answers those commands:

In every world-model setting an execution command is still run in the pod (its side effects persist) and only the observation the agent sees is replaced by the world model's answer; reads, searches and edits are always real. Same settings as the 9/07 SWE-bench page.

Interpreter baseline — every command runs; the agent sees real output.
agentHaiku 4.5bash commandinterpreterruns everything, 60 s capreal output + return code
No code execution (floor) — reads and edits are real; runs return nothing.
agentHaiku 4.5routerrule: does it run code?reads / edits / grepinterpreterreal output + rcreal outputcode executionnothing backoutput withheld · rc hiddenthe run still happens; files it writes are quarantined
All-CWM (original) — the CWM answers everything, using the same retrieved rubrics; the agent cannot read code.
agentHaiku 4.5every command, reads includedCWMgrades the diff vs the retrieved rubricsrubric verdict onlyrubric library10 / 100 / 1000 mined rubricstop-12 nearest to the taskthe agent never sees file contentsor any command's output
Hybrid CWM (this work) — reads go to the interpreter, runs go to the CWM with retrieved rubrics.
agentHaiku 4.5routerrule: does it run code?reads / edits / grepinterpreterreal output + rcreal outputcode executionCWMgrades the diff vs the retrieved rubricsrubric library1000 mined rubricstop-12 nearest totask + diff + commandverdict + real rcthe run executes for side effects only; its output is never shown

The router is a fixed rule, not a model. The command string is split on every bash separator (\n ; && || | &), heredoc bodies are dropped and wrapper prefixes (sudo, timeout N, VAR=x) skipped; a command is code execution if any segment's first word is an interpreter or test runner (python, pytest, tox, make, manage.py, runtests.py, bash, sh, …), a script or repo executable (./run.sh, /testbed/bin/x), or contains a substitution / xargs / find -exec that runs one. Everything else — cat, ls, grep, sed -i, git, heredoc writes — is a read or edit and goes to the interpreter. A mixed command such as cat > t.py <<EOF … EOF && python t.py counts as execution.

The CWM (Haiku 4.5 or Opus 5) reads the task, the agent's current diff and the command, grades the diff against the 12 rubrics retrieved from the mined library (nearest to the task, the files the diff touches and the command; in the hybrid arms a standing task criterion — “does the change implement what the issue asks?” — is added to the list) and returns the per-rubric verdict. Files written by a run are quarantined so the agent cannot read the output back; every trajectory is audited for leakage. No setting ever sees the hidden grading tests.

2Results

any medal, % of 22

0%12%24%36%48%60%9%2/229%2/225%1/225%1/225%1/225%1/229%2/229%2/225%1/225%1/22interpreterrealexecutionno code exec.bashonlyhybrid CWMOpusall + preamblehybrid CWMHaikuall + preambledual 30 sOpusdual 30 sHaikudual 60 sOpusdual 60 sHaikudual 120 sOpusdual 120 sHaiku

above the leaderboard median, % of 22

0%12%24%36%48%60%14%3/229%2/2214%3/225%1/2218%4/229%2/2214%3/2214%3/229%2/229%2/22interpreterrealexecutionno code exec.bashonlyhybrid CWMOpusall + preamblehybrid CWMHaikuall + preambledual 30 sOpusdual 30 sHaikudual 60 sOpusdual 60 sHaikudual 120 sOpusdual 120 sHaiku

MLE-bench's headline metric is the medal rate (top chart); "above the leaderboard median" is the secondary, easier bar. Valid-submission counts are in the table.

armcountratealso
interpreter · real · execution2/229.1%above median 3 · valid 18
no code exec. · bash · only2/229.1%above median 2 · valid 7
hybrid CWM · Opus · all + preamble1/224.5%above median 3 · valid 7
hybrid CWM · Haiku · all + preamble1/224.5%above median 1 · valid 3
dual 30 s · Opus1/224.5%above median 4 · valid 20
dual 30 s · Haiku1/224.5%above median 2 · valid 19
dual 60 s · Opus2/229.1%above median 3 · valid 19
dual 60 s · Haiku2/229.1%above median 3 · valid 19
dual 120 s · Opus1/224.5%above median 2 · valid 19
dual 120 s · Haiku1/224.5%above median 2 · valid 19

Efficiency: median wall-clock per instance

0.0 min26.0 min52.0 min78.0 min104.0 min130.0 min44.8 min20.1 min16.9 min30.3 min29.2 min26.3 min25.7 min32.1 min28.0 min25.2 mininterpreterrealexecutionno code exec.bashonlyhybrid CWMOpusall + preamblehybrid CWMHaikuall + preambledual 30 sOpusdual 30 sHaikudual 60 sOpusdual 60 sHaikudual 120 sOpusdual 120 sHaiku

Wall-clock per instance from the first command to the last (minutes), from the per-command timing records of each pod.

Efficiency: median agent LLM calls per instance

060120180240300581048486575250584546interpreterrealexecutionno code exec.bashonlyhybrid CWMOpusall + preamblehybrid CWMHaikuall + preambledual 30 sOpusdual 30 sHaikudual 60 sOpusdual 60 sHaikudual 120 sOpusdual 120 sHaiku

Efficiency: median seconds per LLM call (wall clock / calls, per instance)

0.0 s10.0 s20.0 s30.0 s40.0 s50.0 s39.8 s10.8 s12.8 s14.0 s19.4 s19.2 s23.6 s21.1 s24.4 s29.5 sinterpreterrealexecutionno code exec.bashonlyhybrid CWMOpusall + preamblehybrid CWMHaikuall + preambledual 30 sOpusdual 30 sHaikudual 60 sOpusdual 60 sHaikudual 120 sOpusdual 120 sHaiku

Wall clock is calls × seconds per call. This chart isolates the second factor: the time the agent waits per step, which contains the model call and, in the world-model arms, the judge verdict; in the interpreter it contains the real execution. An arm that is slower here costs more per step; an arm that is slower only in wall clock took more steps.

Efficiency: median cost per instance (agent + world model, $; judge cost = graded observations x that arm's $/grade)

$0.00$1.20$2.40$3.60$4.80$6.00$0.49$0.53$2.60$0.75$4.24$0.67$3.22$0.69$3.51$0.61interpreterrealexecutionno code exec.bashonlyhybrid CWMOpusall + preamblehybrid CWMHaikuall + preambledual 30 sOpusdual 30 sHaikudual 60 sOpusdual 60 sHaikudual 120 sOpusdual 120 sHaiku

Mock: median wall clock per competition under the canonical architecture (code execution never runs)

0.0 min16.0 min32.0 min48.0 min64.0 min80.0 min54.2 min68.4 min35.7 min9.6 mincanonical mock8.0 minmock9.9 minmock11.9 minmockinterpreteras runno code exec. flooras runhybrid Opus allrun-and-hideas runhybrid Opus allexecution removedinterpreterexec cmds → judge10 sinterpreterexec cmds → judge15 sinterpreterexec cmds → judge20 s

Computed from the recorded per-command timings of the runs above. "Execution removed" subtracts the time of every command the router classified as code execution from the hybrid arm's wall clock (its judge calls are already inside that wall clock). The interpreter rows replace each of its execution commands by one judge call of 10 / 15 / 20 s. Totals over the set: interpreter 23.6 h, hybrid as run 18.0 h, hybrid with execution removed 3.9 h (83% below the interpreter). These are latency mocks: without execution nothing writes the submission file, and agents behave differently when they never see a training result.

What the agent does: command mix per setting

Every agent command classified by the routing rule (inner ring: interpreter vs code execution) and by type (outer ring). "Write file" = heredoc / echo redirects; "edit / file ops" = sed -i, mv, cp…; "run tests" = pytest and friends; "python script" = python x.py; "inline python" = python -c / heredoc.

interpreter — every command real

pending

no code execution — runs return nothing

pending

hybrid CWM · Opus — runs answered by the Opus CWM

pending

dual 60 s · Opus — runs under 60 s + verdict

pending

3Example traces

Trace pair pending: needs the same instance in h2-mle-dual60-opus and h2-mle-cwm-opus.

4Curation cost per rubric

On-policy mining from the DABench + DA-Code data-science pool (run_da_pool.py): the same Haiku agent runs on pool tasks with real execution, every failed rollout is distilled into one lesson by Opus, lessons are deduplicated and cut into the 10 / 100 / all libraries. No CWM grading during mining and no precision filter, so the per-rubric cost is the "cheap" recipe of the SWE-bench page. Prices: Opus 5 $5 / $25 per M tokens, Haiku 4.5 $1 / $5, GKE Spot ≈ $0.045 per node-hour.

itemvalue
pool rollouts (Haiku interpreter agent, real execution)pending
median agent cost per rolloutpending
failures distilled (Opus, one call per failure, ~$0.09 each)pending
lessons after dedup at 0.90 = library sizepending
cost per kept rubric (agent + distillation + pods)pending