← back to the main dashboard · curation cost per rubric

MLE-bench Lite — canonical hybrid CWM

MLE-bench Lite, all 21 competitions held out · Haiku 4.5 agent · official mlebench grader · 2-hour budget per competition · rubrics collected from external data-science task pools only (270 pool rollouts, 69 failures, 69 lessons: the library is 10 / all = 69, and the "100" arms are replicates of "all") · canonical architecture: routed commands are never run (hybrid: code execution; all-CWM: everything); the agent receives only the parsed rubric → reward verdicts as markdown

1Experiment setting

Each task is a Kaggle competition: the agent gets the competition description and data and must leave a valid submission file; the official grader scores it against the real leaderboard. Medal = the submission clears the bronze threshold; the table also reports the above-median rate and valid submissions.

A Haiku 4.5 agent works in the repository by issuing bash commands. The settings differ only in what answers those commands:

Canonical implementation (these are real runs). Same four settings and the same router as the 9/07 page. The one difference from the 9/07 runs: a command the router sends to the world model is never run — earlier runs executed it and hid its output. Reads, searches and edits still run in the pod; the world model grades the real git diff against the retrieved rubrics and the agent receives only the parsed verdicts (reward, violated criteria, checked criteria), nothing else.

Interpreter baseline — every command runs; the agent sees real output.
agentHaiku 4.5bash commandinterpreterruns everything, 60 s capreal output + return code
No code execution (floor) — reads and edits are real; runs return nothing.
agentHaiku 4.5routerrule: does it run code?reads / edits / grepinterpreterreal output + rcreal outputcode executionnothing backoutput withheld · rc hiddenthe command is never run
All-CWM (original) — the CWM answers everything, using the same retrieved rubrics; the agent cannot read code.
agentHaiku 4.5every command, reads includedCWMgrades the diff vs the retrieved rubricsrubric verdict onlyrubric library10 / 100 / 1000 mined rubricstop-12 nearest to the taskthe agent never sees file contentsor any command's output
Hybrid CWM (this work) — reads go to the interpreter, runs go to the CWM with retrieved rubrics.
agentHaiku 4.5routerrule: does it run code?reads / edits / grepinterpreterreal output + rcreal outputcode executionCWMgrades the diff vs the retrieved rubricsrubric library1000 mined rubricstop-12 nearest totask + diff + commandverdict onlythe command is never run; only the verdict comes back

The router is a fixed rule, not a model. The command string is split on every bash separator (\n ; && || | &), heredoc bodies are dropped and wrapper prefixes (sudo, timeout N, VAR=x) skipped; a command is code execution if any segment's first word is an interpreter or test runner (python, pytest, tox, make, manage.py, runtests.py, bash, sh, …), a script or repo executable (./run.sh, /testbed/bin/x), or contains a substitution / xargs / find -exec that runs one. Everything else — cat, ls, grep, sed -i, git, heredoc writes — is a read or edit and goes to the interpreter. A mixed command such as cat > t.py <<EOF … EOF && python t.py counts as execution.

The CWM (Haiku 4.5 or Opus 5) reads the task, the agent's current diff and the command, grades the diff against the 12 rubrics retrieved from the mined library (nearest to the task, the files the diff touches and the command; in the hybrid arms a standing task criterion — “does the change implement what the issue asks?” — is added to the list) and returns the per-rubric verdict. A routed command never runs, so there is no output to leak; every trajectory is audited for leakage and the router is red-teamed (194 cases). No setting ever sees the hidden grading tests.

2Results

any medal, % of 22

0%12%24%36%48%60%14%3/225%1/220%0/220%0/220%0/220%0/220%0/220%0/210%0/220%0/220%0/220%0/220%0/220%0/220%0/220%0/22interpreterrealexecutioninterpreterrealreplicate (09-09)no code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

above the leaderboard median, % of 22

0%12%24%36%48%60%32%7/229%2/220%0/220%0/220%0/220%0/220%0/220%0/215%1/220%0/220%0/220%0/220%0/220%0/220%0/220%0/22interpreterrealexecutioninterpreterrealreplicate (09-09)no code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

MLE-bench's headline metric is the medal rate (top chart); "above the leaderboard median" is the secondary, easier bar. Valid-submission counts are in the table.

above the leaderboard median, % of 22

0%12%24%36%48%60%32%7/229%2/220%0/220%0/220%0/220%0/220%0/220%0/215%1/220%0/220%0/220%0/220%0/220%0/220%0/220%0/22interpreterrealexecutioninterpreterrealreplicate (09-09)no code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc
armcountratealso
interpreter · real · execution3/2213.6%above median 7 · valid 18
interpreter · real · replicate (09-09)1/224.5%above median 2 · valid 18
no code exec. · bash · only0/220.0%above median 0 · valid 19
all-CWM · Opus · 10 rubrics0/220.0%above median 0 · valid 0
all-CWM · Opus · 100 (69-rubric library)0/220.0%above median 0 · valid 0
all-CWM · Opus · all (69-rubric library)0/220.0%above median 0 · valid 0
hybrid CWM · Haiku · task only0/220.0%above median 0 · valid 18
hybrid CWM · Haiku · all rubrics0/210.0%above median 0 · valid 18
hybrid CWM · Opus · task only0/220.0%above median 1 · valid 18
hybrid CWM · Opus · 10 rubrics0/220.0%above median 0 · valid 19
hybrid CWM · Opus · 100 rubrics0/220.0%above median 0 · valid 19
hybrid CWM · Opus · 1000 rubrics0/220.0%above median 0 · valid 19
hybrid CWM · Opus · all rubrics (69-rubric library)0/220.0%above median 0 · valid 19
hybrid CWM · Opus · 12 random (69-rubric library)0/220.0%above median 0 · valid 20
simulator · Haiku · pure0/220.0%above median 0 · valid 14
simulator · Haiku · + real rc0/220.0%above median 0 · valid 6
hybrid CWM · Opus · 6 generated0/220.0%above median 0 · valid 16

Efficiency: median wall-clock per instance

0.0 min26.0 min52.0 min78.0 min104.0 min130.0 min52.5 min46.5 min8.6 min4.8 min4.7 min5.2 min9.9 min9.6 min7.9 min8.6 min11.6 min9.7 min9.4 min13.4 min21.8 min35.5 mininterpreterrealexecutioninterpreterrealreplicate (09-09)no code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

Wall-clock per instance from the first command to the last (minutes), from the per-command timing records of each pod.

Efficiency: median agent LLM calls per instance

060120180240300464812290899511811511011212211411410613494interpreterrealexecutioninterpreterrealreplicate (09-09)no code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

Efficiency: median seconds per LLM call (wall clock / calls, per instance)

0.0 s14.0 s28.0 s42.0 s56.0 s70.0 s51.4 s57.8 s4.8 s3.4 s3.0 s3.3 s5.2 s5.4 s3.8 s4.4 s6.0 s5.0 s5.1 s6.7 s10.1 s18.2 sinterpreterrealexecutioninterpreterrealreplicate (09-09)no code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

Wall clock is calls × seconds per call. This chart isolates the second factor: the time the agent waits per step, which contains the model call and, in the world-model arms, the judge verdict; in the interpreter it contains the real execution. An arm that is slower here costs more per step; an arm that is slower only in wall clock took more steps.

Efficiency: median cost per instance (agent + world model, $; judge cost = graded observations x that arm's $/grade)

$0.00$1.20$2.40$3.60$4.80$6.00$0.38$0.38$0.48$0.40$0.37$0.42$0.50$0.55$0.68$2.08$1.38$0.67$0.59$0.56$0.76$0.66interpreterrealexecutioninterpreterrealreplicate (09-09)no code exec.bashonlyall-CWMOpus10 rubricsall-CWMOpus100 (69-rubric library)all-CWMOpusall (69-rubric library)hybrid CWMHaikutask onlyhybrid CWMHaikuall rubricshybrid CWMOpustask onlyhybrid CWMOpus10 rubricshybrid CWMOpus100 rubricshybrid CWMOpus1000 rubricshybrid CWMOpusall rubrics (69-rubric library)hybrid CWMOpus12 random (69-rubric library)simulatorHaikupuresimulatorHaiku+ real rc

What the agent does: command mix per setting

Every agent command classified by the routing rule (inner ring: interpreter vs code execution) and by type (outer ring). "Write file" = heredoc / echo redirects; "edit / file ops" = sed -i, mv, cp…; "run tests" = pytest and friends; "python script" = python x.py; "inline python" = python -c / heredoc.

interpreter — every command real

pending

no code execution — runs return nothing

pending

hybrid CWM · Opus · all — runs answered by the Opus CWM

pending

hybrid CWM · Haiku · all — runs answered by the Haiku CWM

pending

3Example traces

Trace pair pending: needs the same instance in mle-cwmall-strict and mle-noexec-strict.

4Curation cost per rubric

On-policy mining from the DABench + DA-Code data-science pool (run_da_pool.py): the same Haiku agent runs on pool tasks with real execution, every failed rollout is distilled into one lesson by Opus, lessons are deduplicated and cut into the 10 / 100 / all libraries. No CWM grading during mining and no precision filter, so the per-rubric cost is the "cheap" recipe of the SWE-bench page. Prices: Opus 5 $5 / $25 per M tokens, Haiku 4.5 $1 / $5, GKE Spot ≈ $0.045 per node-hour.

itemvalue
pool rollouts (Haiku interpreter agent, real execution)pending
median agent cost per rolloutpending
failures distilled (Opus, one call per failure, ~$0.09 each)pending
lessons after dedup at 0.90 = library sizepending
cost per kept rubric (agent + distillation + pods)pending