MLE-bench Lite, all 21 competitions held out · Haiku 4.5 agent · official mlebench grader · 2-hour budget per competition · rubrics collected from external data-science task pools only (270 pool rollouts, 69 failures, 69 lessons: the library is 10 / all = 69, and the "100" arms are replicates of "all")
Superseded. This page reports the 9/07 run-and-hide design: execution commands were still run in the pod and only their output was hidden, so wall clock here includes real training/test execution, the side-channel scans after every execution, and (for some arms) API throttling. The canonical never-execute results are on MLE-bench Lite — canonical hybrid CWM.
1Experiment setting
Each task is a Kaggle competition: the agent gets the competition description and data and must leave a valid submission file; the official grader scores it against the real leaderboard. Medal = the submission clears the bronze threshold; the table also reports the above-median rate and valid submissions.
A Haiku 4.5 agent works in the repository by issuing bash commands. The settings differ only in what answers those commands:
In every world-model setting an execution command is still run in the pod (its side effects persist) and only the observation the agent sees is replaced by the world model's answer; reads, searches and edits are always real. Same settings as the 9/07 SWE-bench page.
Interpreter baseline — every command runs; the agent sees real output.No code execution (floor) — reads and edits are real; runs return nothing.All-CWM (original) — the CWM answers everything, using the same retrieved rubrics; the agent cannot read code.Hybrid CWM (this work) — reads go to the interpreter, runs go to the CWM with retrieved rubrics.
The router is a fixed rule, not a model. The command string is split on every bash separator (\n ; && || | &), heredoc bodies are dropped and wrapper prefixes (sudo, timeout N, VAR=x) skipped; a command is code execution if any segment's first word is an interpreter or test runner (python, pytest, tox, make, manage.py, runtests.py, bash, sh, …), a script or repo executable (./run.sh, /testbed/bin/x), or contains a substitution / xargs / find -exec that runs one. Everything else — cat, ls, grep, sed -i, git, heredoc writes — is a read or edit and goes to the interpreter. A mixed command such as cat > t.py <<EOF … EOF && python t.py counts as execution.
The CWM (Haiku 4.5 or Opus 5) reads the task, the agent's current diff and the command, grades the diff against the 12 rubrics retrieved from the mined library (nearest to the task, the files the diff touches and the command; in the hybrid arms a standing task criterion — “does the change implement what the issue asks?” — is added to the list) and returns the per-rubric verdict. Files written by a run are quarantined so the agent cannot read the output back; every trajectory is audited for leakage. No setting ever sees the hidden grading tests.
2Results
any medal, % of 22
above the leaderboard median, % of 22
MLE-bench's headline metric is the medal rate (top chart); "above the leaderboard median" is the secondary, easier bar. Valid-submission counts are in the table.
above the leaderboard median, % of 22
arm
count
rate
also
interpreter · real · execution
3/22
13.6%
above median 7 · valid 18
no code exec. · bash · only
0/22
0.0%
above median 2 · valid 16
all-CWM · Opus · 10 rubrics
0/22
0.0%
above median 0 · valid 0
all-CWM · Opus · 100 (69-rubric library)
0/22
0.0%
above median 0 · valid 0
all-CWM · Opus · all (69-rubric library)
0/21
0.0%
above median 0 · valid 0
hybrid CWM · Haiku · task only
0/22
0.0%
above median 1 · valid 16
hybrid CWM · Haiku · all rubrics
0/21
0.0%
above median 1 · valid 12
hybrid CWM · Opus · task only
0/22
0.0%
above median 1 · valid 15
hybrid CWM · Opus · 10 rubrics
0/22
0.0%
above median 3 · valid 19
hybrid CWM · Opus · 100 rubrics
0/21
0.0%
above median 1 · valid 14
hybrid CWM · Opus · 1000 rubrics
pending
hybrid CWM · Opus · all rubrics (69-rubric library)
1/22
4.5%
above median 1 · valid 18
hybrid CWM · Opus · 12 random (69-rubric library)
1/21
4.8%
above median 3 · valid 13
simulator · Haiku · pure
0/22
0.0%
above median 1 · valid 9
simulator · Haiku · + real rc
0/22
0.0%
above median 0 · valid 6
hybrid CWM · Opus · 6 generated
1/22
4.5%
above median 3 · valid 16
Efficiency: median wall-clock per instance
Wall-clock per instance from the first command to the last (minutes), from the per-command timing records of each pod.
Efficiency: median agent LLM calls per instance
Efficiency: median seconds per LLM call (wall clock / calls, per instance)
Wall clock is calls × seconds per call. This chart isolates the second factor: the time the agent waits per step, which contains the model call and, in the world-model arms, the judge verdict; in the interpreter it contains the real execution. An arm that is slower here costs more per step; an arm that is slower only in wall clock took more steps.
Efficiency: median cost per instance (agent + world model, $; judge cost = graded observations x that arm's $/grade)
Mock: median wall clock per competition under the canonical architecture (code execution never runs)
Computed from the recorded per-command timings of the runs above. "Execution removed" subtracts the time of every command the router classified as code execution from the hybrid arm's wall clock (its judge calls are already inside that wall clock). The interpreter rows replace each of its execution commands by one judge call of 10 / 15 / 20 s. Totals over the set: interpreter 23.6 h, hybrid as run 18.0 h, hybrid with execution removed 3.9 h (83% below the interpreter). These are latency mocks: without execution nothing writes the submission file, and agents behave differently when they never see a training result.
What the agent does: command mix per setting
Every agent command classified by the routing rule (inner ring: interpreter vs code execution) and by type (outer ring). "Write file" = heredoc / echo redirects; "edit / file ops" = sed -i, mv, cp…; "run tests" = pytest and friends; "python script" = python x.py; "inline python" = python -c / heredoc.
interpreter — every command real
pending
no code execution — runs return nothing
pending
hybrid CWM · Opus · all — runs answered by the Opus CWM
pending
hybrid CWM · Haiku · all — runs answered by the Haiku CWM
pending
3Example traces
Trace pair pending: needs the same instance in mle-cwmall and mle-noexec.
4Curation cost per rubric
On-policy mining from the DABench + DA-Code data-science pool (run_da_pool.py): the same Haiku agent runs on pool tasks with real execution, every failed rollout is distilled into one lesson by Opus, lessons are deduplicated and cut into the 10 / 100 / all libraries. No CWM grading during mining and no precision filter, so the per-rubric cost is the "cheap" recipe of the SWE-bench page. Prices: Opus 5 $5 / $25 per M tokens, Haiku 4.5 $1 / $5, GKE Spot ≈ $0.045 per node-hour.
item
value
pool rollouts (Haiku interpreter agent, real execution)
pending
median agent cost per rollout
pending
failures distilled (Opus, one call per failure, ~$0.09 each)
pending
lessons after dedup at 0.90 = library size
pending
cost per kept rubric (agent + distillation + pods)