← back to the main dashboard · curation cost per rubric
MLE-bench Lite, all 21 competitions held out · Haiku 4.5 agent · official mlebench grader · 2-hour budget per competition · rubrics collected from external data-science task pools only (270 pool rollouts, 69 failures, 69 lessons: the library is 10 / all = 69, and the "100" arms are replicates of "all") · canonical architecture: routed commands are never run (hybrid: code execution; all-CWM: everything); the agent receives only the parsed rubric → reward verdicts as markdown
Each task is a Kaggle competition: the agent gets the competition description and data and must leave a valid submission file; the official grader scores it against the real leaderboard. Medal = the submission clears the bronze threshold; the table also reports the above-median rate and valid submissions.
A Haiku 4.5 agent works in the repository by issuing bash commands. The settings differ only in what answers those commands:
Canonical implementation (these are real runs). Same four settings and the same router as the 9/07 page. The one difference from the 9/07 runs: a command the router sends to the world model is never run — earlier runs executed it and hid its output. Reads, searches and edits still run in the pod; the world model grades the real git diff against the retrieved rubrics and the agent receives only the parsed verdicts (reward, violated criteria, checked criteria), nothing else.
The router is a fixed rule, not a model. The command string is split on every bash separator (\n ; && || | &), heredoc bodies are dropped and wrapper prefixes (sudo, timeout N, VAR=x) skipped; a command is code execution if any segment's first word is an interpreter or test runner (python, pytest, tox, make, manage.py, runtests.py, bash, sh, …), a script or repo executable (./run.sh, /testbed/bin/x), or contains a substitution / xargs / find -exec that runs one. Everything else — cat, ls, grep, sed -i, git, heredoc writes — is a read or edit and goes to the interpreter. A mixed command such as cat > t.py <<EOF … EOF && python t.py counts as execution.
The CWM (Haiku 4.5 or Opus 5) reads the task, the agent's current diff and the command, grades the diff against the 12 rubrics retrieved from the mined library (nearest to the task, the files the diff touches and the command; in the hybrid arms a standing task criterion — “does the change implement what the issue asks?” — is added to the list) and returns the per-rubric verdict. A routed command never runs, so there is no output to leak; every trajectory is audited for leakage and the router is red-teamed (194 cases). No setting ever sees the hidden grading tests.
MLE-bench's headline metric is the medal rate (top chart); "above the leaderboard median" is the secondary, easier bar. Valid-submission counts are in the table.
| arm | count | rate | also |
|---|---|---|---|
| interpreter · real · execution | 3/22 | 13.6% | above median 7 · valid 18 |
| interpreter · real · replicate (09-09) | 1/22 | 4.5% | above median 2 · valid 18 |
| no code exec. · bash · only | 0/22 | 0.0% | above median 0 · valid 19 |
| all-CWM · Opus · 10 rubrics | 0/22 | 0.0% | above median 0 · valid 0 |
| all-CWM · Opus · 100 (69-rubric library) | 0/22 | 0.0% | above median 0 · valid 0 |
| all-CWM · Opus · all (69-rubric library) | 0/22 | 0.0% | above median 0 · valid 0 |
| hybrid CWM · Haiku · task only | 0/22 | 0.0% | above median 0 · valid 18 |
| hybrid CWM · Haiku · all rubrics | 0/21 | 0.0% | above median 0 · valid 18 |
| hybrid CWM · Opus · task only | 0/22 | 0.0% | above median 1 · valid 18 |
| hybrid CWM · Opus · 10 rubrics | 0/22 | 0.0% | above median 0 · valid 19 |
| hybrid CWM · Opus · 100 rubrics | 0/22 | 0.0% | above median 0 · valid 19 |
| hybrid CWM · Opus · 1000 rubrics | 0/22 | 0.0% | above median 0 · valid 19 |
| hybrid CWM · Opus · all rubrics (69-rubric library) | 0/22 | 0.0% | above median 0 · valid 19 |
| hybrid CWM · Opus · 12 random (69-rubric library) | 0/22 | 0.0% | above median 0 · valid 20 |
| simulator · Haiku · pure | 0/22 | 0.0% | above median 0 · valid 14 |
| simulator · Haiku · + real rc | 0/22 | 0.0% | above median 0 · valid 6 |
| hybrid CWM · Opus · 6 generated | 0/22 | 0.0% | above median 0 · valid 16 |
Wall-clock per instance from the first command to the last (minutes), from the per-command timing records of each pod.
Wall clock is calls × seconds per call. This chart isolates the second factor: the time the agent waits per step, which contains the model call and, in the world-model arms, the judge verdict; in the interpreter it contains the real execution. An arm that is slower here costs more per step; an arm that is slower only in wall clock took more steps.
Every agent command classified by the routing rule (inner ring: interpreter vs code execution) and by type (outer ring). "Write file" = heredoc / echo redirects; "edit / file ops" = sed -i, mv, cp…; "run tests" = pytest and friends; "python script" = python x.py; "inline python" = python -c / heredoc.
pending
pending
pending
pending
Trace pair pending: needs the same instance in mle-cwmall-strict and mle-noexec-strict.
On-policy mining from the DABench + DA-Code data-science pool (run_da_pool.py): the same Haiku agent runs on pool tasks with real execution, every failed rollout is distilled into one lesson by Opus, lessons are deduplicated and cut into the 10 / 100 / all libraries. No CWM grading during mining and no precision filter, so the per-rubric cost is the "cheap" recipe of the SWE-bench page. Prices: Opus 5 $5 / $25 per M tokens, Haiku 4.5 $1 / $5, GKE Spot ≈ $0.045 per node-hour.
| item | value |
|---|---|
| pool rollouts (Haiku interpreter agent, real execution) | pending |
| median agent cost per rollout | pending |
| failures distilled (Opus, one call per failure, ~$0.09 each) | pending |
| lessons after dedup at 0.90 = library size | pending |
| cost per kept rubric (agent + distillation + pods) | pending |