Hybrid CWM: efficiency on solved instances, and what the rubrics cover
SWE-bench Verified and MLE-bench Lite · efficiency measured only where the method actually solved the task · test-time errors vs rubric coverage in one 20-bucket taxonomy · built 2026-09-16 09:21
Why filter. Unfiltered efficiency numbers reward arms that give up: an agent that submits a wrong patch after five commands has a short wall clock and a low bill. Every efficiency statistic on this page is therefore reported on the instances the arm solved, and, for the fairest reading, on the instances both the arm and the interpreter solved, so the same problems are being compared.
SWE-bench Verified (500), Haiku 4.5 agent
Setup. Per instance, wall clock = first agent command to the end of the last one, from timing.json. Sandbox time = summed duration of commands that reached the sandbox (in hybrid arms a routed command only contributes its git diff). Execution / judge time = real execution time plus the judge's latency; judge latency is measured from the timing records where the harness recorded it (harness 2 arms) and estimated otherwise (see the note under that chart). Agent time = the rest (LLM calls and harness). Calls and $ from the run logs; judge $ = verdict count × that arm's logged $ per grade.
Filters.all: every instance the arm ran. solved by the arm: the official FAIL_TO_PASS / PASS_TO_PASS grade. solved by both: the arm and the interpreter both solved it, so the same problems are compared. An arm that breaks early on hard problems looks fast on 'all' and loses that advantage under the paired filter.
Median wall clock per instance (minutes) under the three filters
Wall clock per instance, solved-by-both instances (p5–p95 whiskers, p25–p75 box, median line)
Time on code-execution commands, solved-by-both: real execution (interpreter) or estimated judge latency (hybrid)
Judge latency is not logged; it is estimated per instance as wall clock minus sandbox time minus the agent's calls x the interpreter's per-call latency (1.9 s per call on its solved instances). Arms ran on different days; Opus overload periods inflate judge latency.
{"exec + judge, paired" if judge_est else "sandbox time, paired"}
judge, paired
agent time, paired
calls, paired
$, paired
routed cmds
verdicts
s per routed cmd
interpreter
339/500
339
3.0
2.6
2.6
0.5
0
1.6
51
$0.19
19
0
1.5
floor: no execution
297/500
269
3.9
3.5
3.4
0.0
0
2.8
87
$0.31
28
0
0.0
hybrid CWM · Opus judge · mined 1000
313/500
281
11.3
10.0
9.7
3.2
3.2 (est.)
2.1
67
$4.01
19
19
10.1
hybrid CWM · Opus judge · mined 1000 + preamble
315/500
280
4.4
3.9
3.7
0.4
0.4 (est.)
2.1
68
$3.62
17
17
1.4
hybrid CWM · Haiku judge · mined 1000
307/500
277
11.6
10.5
10.2
3.7
3.7 (est.)
2.2
68
$0.56
20
20
11.0
Minutes are medians. 'routed cmds' = median number of code-execution commands per instance (router class); 'verdicts' = world-model replies; 's per routed cmd' = execution (interpreter) or estimated judge latency (hybrid) per such command. Arms ran on different days, so judge latency also reflects the API load at the time (e.g. Opus overload periods); compare cost and call counts, which do not depend on that.
What this says. floor: no execution: wall 3.4 vs interpreter 2.6 min on the same 269 solved instances (+34%); exec+judge 0.0 (judge 0.0) vs execution 0.5 min; calls 87 vs 51; cost $0.31 vs $0.19 hybrid CWM · Opus judge · mined 1000: wall 9.7 vs interpreter 2.6 min on the same 281 solved instances (+278%); exec+judge 3.2 (judge 3.2 est.) vs execution 0.5 min; calls 67 vs 51; cost $4.01 vs $0.19 hybrid CWM · Opus judge · mined 1000 + preamble: wall 3.7 vs interpreter 2.6 min on the same 280 solved instances (+46%); exec+judge 0.4 (judge 0.4 est.) vs execution 0.5 min; calls 68 vs 51; cost $3.62 vs $0.19 hybrid CWM · Haiku judge · mined 1000: wall 10.2 vs interpreter 2.6 min on the same 277 solved instances (+300%); exec+judge 3.7 (judge 3.7 est.) vs execution 0.5 min; calls 68 vs 51; cost $0.56 vs $0.19
MLE-bench Lite (22), Haiku 4.5 agent, solved = beats the no-execution floor
Setup. Per instance, wall clock = first agent command to the end of the last one, from timing.json (MLE data-setup commands excluded). Sandbox time = summed duration of commands that reached the sandbox (in hybrid arms a routed command only contributes its git diff). Execution / judge time = real execution time plus the judge's latency; judge latency is measured from the timing records where the harness recorded it (harness 2 arms) and estimated otherwise (see the note under that chart). Agent time = the rest (LLM calls and harness). Calls and $ from the run logs; judge $ = verdict count × that arm's logged $ per grade.
Filters.all: every instance the arm ran. solved by the arm: a valid submission from a run that finished (exit status 'submitted') whose score, relative to the leaderboard median, is better than the no-execution floor's on the same competition by at least 0.02. This excludes constant-prediction files (valid, but identical to doing nothing) without demanding a medal. The threshold sweep below shows how every count depends on the bar. solved by both: the arm and the interpreter both solved it, so the same problems are compared. An arm that breaks early on hard problems looks fast on 'all' and loses that advantage under the paired filter.
Median wall clock per instance (minutes) under the three filters
Paired n: interpreter: 11, floor: no execution: 0, hybrid CWM · Opus judge · all rubrics: 5, hybrid CWM · Opus judge · 1000 rubrics: 6, hybrid CWM · Haiku judge · all rubrics: 4.
Wall clock per instance, solved-by-both instances (p5–p95 whiskers, p25–p75 box, median line)
Sandbox time (recorded command durations), solved-by-both instances
On MLE-bench the interpreter's remaining time per call is dominated by long data-loading and training waits, so the judge-latency estimate used on SWE-bench is not meaningful here; only recorded sandbox time is shown.
Agent LLM calls per instance, solved-by-both
Cost per instance, solved-by-both
arm
solved
paired n
wall, all
wall, solved
wall, paired
{"exec + judge, paired" if judge_est else "sandbox time, paired"}
judge, paired
agent time, paired
calls, paired
$, paired
routed cmds
verdicts
s per routed cmd
interpreter
11/22
11
46.5
42.3
42.3
38.9
0
n/a
42
$0.37
23
0
100.5
floor: no execution
0/22
0
8.6
n/a
n/a
n/a
0
n/a
n/a
n/a
n/a
n/a
n/a
hybrid CWM · Opus judge · all rubrics
7/22
5
9.4
7.3
7.1
2.3
0
n/a
114
$8.10
45
45
0.0
hybrid CWM · Opus judge · 1000 rubrics
9/22
6
9.7
9.1
9.8
5.4
0
n/a
114
$8.21
34
34
0.0
hybrid CWM · Haiku judge · all rubrics
6/21
4
9.6
7.0
6.7
2.7
0
n/a
108
$0.95
30
30
0.0
Minutes are medians. 'routed cmds' = median number of code-execution commands per instance (router class); 'verdicts' = world-model replies; 's per routed cmd' = execution (interpreter) or estimated judge latency (hybrid) per such command. Arms ran on different days, so judge latency also reflects the API load at the time (e.g. Opus overload periods); compare cost and call counts, which do not depend on that.
Per arm: beats floor / >= 90% of median / above median / medal / valid: interpreter: 11 / 7 / 2 / 1 / 18, floor: no execution: 0 / 1 / 0 / 0 / 19, hybrid CWM · Opus judge · all rubrics: 7 / 3 / 0 / 0 / 19, hybrid CWM · Opus judge · 1000 rubrics: 9 / 2 / 0 / 0 / 19, hybrid CWM · Haiku judge · all rubrics: 6 / 0 / 0 / 0 / 18.
MLE-bench: every valid submission, score relative to the leaderboard median vs wall clock
One point per competition and arm (hover for the name). Points on the same horizontal line across arms are constant-prediction submissions that score like the floor. The hybrid's speed comes from never training a model; only points above the competitive bar count as solved.
Threshold sweep: competitions counted as solved at each bar (finished, valid, score/median >= bar)
At 50% of the median even the floor passes most competitions; at the median the hybrids pass none.
bar
interpreter solved
floor: no execution: paired n · wall hybrid vs interpreter
hybrid CWM · Opus judge · all rubrics: paired n · wall hybrid vs interpreter
hybrid CWM · Opus judge · 1000 rubrics: paired n · wall hybrid vs interpreter
hybrid CWM · Haiku judge · all rubrics: paired n · wall hybrid vs interpreter
50% of median
14
11 · 8 vs 41 min
10 · 9 vs 41 min
13 · 9 vs 49 min
10 · 8 vs 45 min
60% of median
11
5 · 11 vs 28 min
4 · 7 vs 27 min
6 · 13 vs 46 min
2 · 9 vs 19 min
70% of median
10
3 · 8 vs 28 min
4 · 7 vs 27 min
5 · 13 vs 41 min
2 · 9 vs 19 min
75% of median
10
3 · 8 vs 28 min
3 · 7 vs 28 min
5 · 13 vs 41 min
2 · 9 vs 19 min
80% of median
8
2 · 12 vs 29 min
3 · 7 vs 28 min
4 · 16 vs 46 min
2 · 9 vs 19 min
85% of median
8
1 · 16 vs 49 min
2 · 8 vs 30 min
2 · 13 vs 50 min
1 · 12 vs 9 min
90% of median
7
1 · 16 vs 49 min
1 · 10 vs 50 min
1 · 5 vs 9 min
0
at the median
2
0
0
0
0
Paired n = competitions both the arm and the interpreter pass at that bar; minutes are medians over those competitions.
What this says. Paired sets are tiny here (floor: no execution: 0, hybrid CWM · Haiku judge · all rubrics: 4 competitions solved by both), so the lines below are illustrative, not a measurement. hybrid CWM · Opus judge · all rubrics: wall 7.1 vs interpreter 42.3 min on the same 5 solved instances (-83%); sandbox time 2.3 vs 38.9 min; calls 114 vs 42; cost $8.10 vs $0.37 hybrid CWM · Opus judge · 1000 rubrics: wall 9.8 vs interpreter 42.3 min on the same 6 solved instances (-77%); sandbox time 5.4 vs 38.9 min; calls 114 vs 42; cost $8.21 vs $0.37 hybrid CWM · Haiku judge · all rubrics: wall 6.7 vs interpreter 42.3 min on the same 4 solved instances (-84%); sandbox time 2.7 vs 38.9 min; calls 108 vs 42; cost $0.95 vs $0.37
Opus vs Haiku as the world model, same library, on the instances each arm solved that the interpreter also solved
benchmark
solved (Opus / Haiku)
paired n
median wall min
median exec / judge min (est.)
median calls
median $
SWE-bench Verified
313 / 307
281 / 277
9.7 / 10.2
3.2 / 3.7
67 / 68
$4.01 / $0.56
MLE-bench Lite (beats the floor; exec/judge column = sandbox time)
7 / 6
5 / 4
7.1 / 6.7
0.0 / 0.0
114 / 108
$8.10 / $0.95
What this says. On SWE-bench, Opus solves 313 vs Haiku 307; on the instances each solved that the interpreter also solved, Opus is faster per instance (10.2 vs 9.7 min) and Haiku is cheaper ($0.56 vs $4.01). Estimated judge time per instance: Opus 3.2 min vs Haiku 3.7. Opus remains the better judge on outcome; Haiku's advantage is only latency and cost, and both hybrids stay behind the interpreter on resolve.
Harness 2 (9/14): interpreter vs CWM-only vs dual channel
What is different. Harness pulled 2026-09-14: the agent is told which channel it is on, package installs never routed, 120 s grader timeout, reopening circuit breaker, and every judge call is timed, so judge time below is measured, not estimated. CWM only = routed commands never run, verdict only (mined 1000 / on-policy all + task criterion + preamble). Dual N s = the routed command runs for real and is killed at N seconds; the agent gets its output (or the partial output and a 'killed' line) and the verdict. MLE arms all write final.py, run once uncapped at collection. Plan: docs/PLAN_H2_0914.md.
SWE-bench Verified (500), harness 2
Setup. Per instance, wall clock = first agent command to the end of the last one, from timing.json. Sandbox time = summed duration of commands that reached the sandbox (in hybrid arms a routed command only contributes its git diff). Execution / judge time = real execution time plus the judge's latency; judge latency is measured from the timing records where the harness recorded it (harness 2 arms) and estimated otherwise (see the note under that chart). Agent time = the rest (LLM calls and harness). Calls and $ from the run logs; judge $ = verdict count × that arm's logged $ per grade.
Filters.all: every instance the arm ran. solved by the arm: the official FAIL_TO_PASS / PASS_TO_PASS grade. solved by both: the arm and the interpreter both solved it, so the same problems are compared. An arm that breaks early on hard problems looks fast on 'all' and loses that advantage under the paired filter.
Median wall clock per instance (minutes) under the three filters
Paired n: interpreter: 347, floor: no execution: 300, CWM only · Opus: 300, CWM only · Haiku: 291, dual 30 s · Opus: 319, dual 30 s · Haiku: 317, dual 60 s · Opus: 320, dual 60 s · Haiku: 315, dual 120 s · Opus: 324, dual 120 s · Haiku: 319.
Wall clock per instance, solved-by-both instances (p5–p95 whiskers, p25–p75 box, median line)
Time on code-execution commands, solved-by-both: real execution (interpreter) or estimated judge latency (hybrid)
Judge latency is not logged; it is estimated per instance as wall clock minus sandbox time minus the agent's calls x the interpreter's per-call latency (3.5 s per call on its solved instances). Arms ran on different days; Opus overload periods inflate judge latency.
{"exec + judge, paired" if judge_est else "sandbox time, paired"}
judge, paired
agent time, paired
calls, paired
$, paired
routed cmds
verdicts
s per routed cmd
interpreter
347/500
347
4.8
4.4
4.4
0.7
0
3.0
53
$0.20
21
0
2.0
floor: no execution
325/500
300
5.1
4.2
4.0
0.0
0
3.0
55
$1.82
8
9
0.0
CWM only · Opus
324/500
300
4.7
3.8
3.7
0.3
0.3 (est.)
2.3
50
$1.46
5
6
4.1
CWM only · Haiku
307/500
291
5.4
4.3
4.3
0.5
0.5 (est.)
2.7
53
$0.33
6
7
4.9
dual 30 s · Opus
346/500
319
6.5
5.1
5.0
1.9
1.4
1.9
40
$2.40
13
11
8.9
dual 30 s · Haiku
340/500
317
6.9
6.0
5.8
2.5
2.1
2.1
42
$0.35
14
12
10.9
dual 60 s · Opus
348/500
320
6.3
5.1
4.9
1.9
1.5
1.8
41
$2.54
13
11
9.0
dual 60 s · Haiku
339/500
315
6.9
5.9
5.8
2.5
2.1
2.0
42
$0.34
13
12
11.7
dual 120 s · Opus
352/500
324
6.3
5.3
5.0
2.0
1.5
1.9
40
$2.58
13
11
9.2
dual 120 s · Haiku
342/500
319
6.7
5.9
5.7
2.6
2.0
2.0
42
$0.34
13
12
11.9
Minutes are medians. 'routed cmds' = median number of code-execution commands per instance (router class); 'verdicts' = world-model replies; 's per routed cmd' = execution (interpreter) or estimated judge latency (hybrid) per such command. Arms ran on different days, so judge latency also reflects the API load at the time (e.g. Opus overload periods); compare cost and call counts, which do not depend on that.
What this says. floor: no execution: wall 4.0 vs interpreter 4.4 min on the same 300 solved instances (-10%); exec+judge 0.0 (judge 0.0 est.) vs execution 0.7 min; calls 55 vs 53; cost $1.82 vs $0.20 CWM only · Opus: wall 3.7 vs interpreter 4.4 min on the same 300 solved instances (-17%); exec+judge 0.3 (judge 0.3 est.) vs execution 0.7 min; calls 50 vs 53; cost $1.46 vs $0.20 CWM only · Haiku: wall 4.3 vs interpreter 4.4 min on the same 291 solved instances (-4%); exec+judge 0.5 (judge 0.5 est.) vs execution 0.7 min; calls 53 vs 53; cost $0.33 vs $0.20 dual 30 s · Opus: wall 5.0 vs interpreter 4.4 min on the same 319 solved instances (+13%); exec+judge 1.9 (judge 1.4) vs execution 0.7 min; calls 40 vs 53; cost $2.40 vs $0.20 dual 30 s · Haiku: wall 5.8 vs interpreter 4.4 min on the same 317 solved instances (+31%); exec+judge 2.5 (judge 2.1) vs execution 0.7 min; calls 42 vs 53; cost $0.35 vs $0.20 dual 60 s · Opus: wall 4.9 vs interpreter 4.4 min on the same 320 solved instances (+11%); exec+judge 1.9 (judge 1.5) vs execution 0.7 min; calls 41 vs 53; cost $2.54 vs $0.20 dual 60 s · Haiku: wall 5.8 vs interpreter 4.4 min on the same 315 solved instances (+31%); exec+judge 2.5 (judge 2.1) vs execution 0.7 min; calls 42 vs 53; cost $0.34 vs $0.20 dual 120 s · Opus: wall 5.0 vs interpreter 4.4 min on the same 324 solved instances (+14%); exec+judge 2.0 (judge 1.5) vs execution 0.7 min; calls 40 vs 53; cost $2.58 vs $0.20 dual 120 s · Haiku: wall 5.7 vs interpreter 4.4 min on the same 319 solved instances (+28%); exec+judge 2.6 (judge 2.0) vs execution 0.7 min; calls 42 vs 53; cost $0.34 vs $0.20
MLE-bench Lite (22), harness 2, solved = beats the harness-2 floor
Setup. Per instance, wall clock = first agent command to the end of the last one, from timing.json (MLE data-setup commands excluded). Sandbox time = summed duration of commands that reached the sandbox (in hybrid arms a routed command only contributes its git diff). Execution / judge time = real execution time plus the judge's latency; judge latency is measured from the timing records where the harness recorded it (harness 2 arms) and estimated otherwise (see the note under that chart). Agent time = the rest (LLM calls and harness). Calls and $ from the run logs; judge $ = verdict count × that arm's logged $ per grade.
Filters.all: every instance the arm ran. solved by the arm: finished, valid, score/median better than h2-mle-floor's on the same competition by 0.02. solved by both: the arm and the interpreter both solved it, so the same problems are compared. An arm that breaks early on hard problems looks fast on 'all' and loses that advantage under the paired filter.
Median wall clock per instance (minutes) under the three filters
Paired n: interpreter: 4, floor: no execution: 0, CWM only · Opus: 1, CWM only · Haiku: 1, dual 30 s · Opus: 0, dual 30 s · Haiku: 1, dual 60 s · Opus: 1, dual 60 s · Haiku: 1, dual 120 s · Opus: 0, dual 120 s · Haiku: 3.
Wall clock per instance, solved-by-both instances (p5–p95 whiskers, p25–p75 box, median line)
Sandbox time (recorded command durations), solved-by-both instances
On MLE-bench the interpreter's remaining time per call is dominated by long data-loading and training waits, so the judge-latency estimate used on SWE-bench is not meaningful here; only recorded sandbox time is shown.
Agent LLM calls per instance, solved-by-both
Cost per instance, solved-by-both
arm
solved
paired n
wall, all
wall, solved
wall, paired
{"exec + judge, paired" if judge_est else "sandbox time, paired"}
judge, paired
agent time, paired
calls, paired
$, paired
routed cmds
verdicts
s per routed cmd
interpreter
4/22
4
44.8
68.7
68.7
60.2
0
n/a
63
$0.58
20
0
182.1
floor: no execution
0/22
0
20.1
n/a
n/a
n/a
0
n/a
n/a
n/a
n/a
n/a
n/a
CWM only · Opus
2/22
1
16.9
30.6
38.0
28.3
2.3
n/a
56
$3.53
13
13
135.1
CWM only · Haiku
1/22
1
30.3
21.5
21.5
7.4
3.7
n/a
68
$0.77
14
14
42.0
dual 30 s · Opus
0/22
0
29.2
n/a
n/a
n/a
0
n/a
n/a
n/a
n/a
n/a
n/a
dual 30 s · Haiku
1/22
1
26.3
27.3
27.3
5.6
3.2
n/a
91
$1.09
16
14
29.9
dual 60 s · Opus
2/22
1
25.7
13.0
17.7
9.0
2.3
n/a
54
$4.49
21
17
30.8
dual 60 s · Haiku
2/22
1
32.1
50.1
85.8
18.3
4.1
n/a
244
$2.54
24
20
49.9
dual 120 s · Opus
0/22
0
28.0
n/a
n/a
n/a
0
n/a
n/a
n/a
n/a
n/a
n/a
dual 120 s · Haiku
3/22
3
25.2
25.9
25.9
11.3
4.3
n/a
38
$0.64
18
16
52.2
Minutes are medians. 'routed cmds' = median number of code-execution commands per instance (router class); 'verdicts' = world-model replies; 's per routed cmd' = execution (interpreter) or estimated judge latency (hybrid) per such command. Arms ran on different days, so judge latency also reflects the API load at the time (e.g. Opus overload periods); compare cost and call counts, which do not depend on that.
What this says. Paired sets are tiny here (floor: no execution: 0, CWM only · Opus: 1, CWM only · Haiku: 1, dual 30 s · Opus: 0, dual 30 s · Haiku: 1, dual 60 s · Opus: 1, dual 60 s · Haiku: 1, dual 120 s · Opus: 0, dual 120 s · Haiku: 3 competitions solved by both), so the lines below are illustrative, not a measurement. CWM only · Opus: wall 38.0 vs interpreter 68.7 min on the same 1 solved instances (-45%); sandbox time 28.3 vs 60.2 min; calls 56 vs 63; cost $3.53 vs $0.58 CWM only · Haiku: wall 21.5 vs interpreter 68.7 min on the same 1 solved instances (-69%); sandbox time 7.4 vs 60.2 min; calls 68 vs 63; cost $0.77 vs $0.58 dual 30 s · Haiku: wall 27.3 vs interpreter 68.7 min on the same 1 solved instances (-60%); sandbox time 5.6 vs 60.2 min; calls 91 vs 63; cost $1.09 vs $0.58 dual 60 s · Opus: wall 17.7 vs interpreter 68.7 min on the same 1 solved instances (-74%); sandbox time 9.0 vs 60.2 min; calls 54 vs 63; cost $4.49 vs $0.58 dual 60 s · Haiku: wall 85.8 vs interpreter 68.7 min on the same 1 solved instances (+25%); sandbox time 18.3 vs 60.2 min; calls 244 vs 63; cost $2.54 vs $0.58 dual 120 s · Haiku: wall 25.9 vs interpreter 68.7 min on the same 3 solved instances (-62%); sandbox time 11.3 vs 60.2 min; calls 38 vs 63; cost $0.64 vs $0.58
Test-time feedback vs rubric coverage, harness-2 traces (buckets re-induced from these traces)
How the buckets were made. Bottom-up from the corpus, not by hand. Corpus = every observation the interpreter agent received from a command the router classes as code execution (harness-2 interpreter traces: SWE-bench h2-interp and MLE-bench h2-mle-interp, successes included) plus every rubric in the libraries the hybrid arms retrieve from (mined 1000 for SWE, on-policy all for MLE). Opus open-coded 48 stratified chunks of 50 items (feedback and rubric chunks, both benchmarks) into candidate categories, then merged all candidates into exactly 20 mutually exclusive buckets with definitions and examples. Haiku then labelled every observation and every rubric with one bucket (rubrics: plus an optional secondary). Code: analysis/taxonomy_induce.py.
Reading the charts. Three bars per bucket: all interpreter feedback (what the agent hears back, successes included), the error-like subset (non-zero return code or an error signature), and the rubrics (by the signal a violation would produce). Mismatch = total variation distance between two distributions in percentage points (0 = identical, 100 = disjoint).
SWE-bench Verified: share per bucket
12,840 feedback observations (3,126 error-like) from 500 instances; 1,000 rubrics. Mismatch rubrics vs all feedback 85, vs error-like feedback 73.
MLE-bench Lite: share per bucket
606 feedback observations (130 error-like) from 22 instances; 1,076 rubrics. Mismatch rubrics vs all feedback 97, vs error-like feedback 98.
The 20 induced buckets
id
bucket
definition
examples
sides
B01
All tests pass
A real test-runner invocation (pytest, unittest, runtests.py) completes with every selected test passing, possibly with skips, xfails or warnings, and no failures or errors.
OK — 42 tests passed, 3 skipped; '1 passed in 0.31s' with no failures reported; Ran 17 tests ... OK (skipped=2)
feedback
B02
Verified successful run or artifact
A script, pipeline, or ad-hoc check runs to completion and produces positive evidence — PASS/SUCCESS markers, training metrics, or an output file whose format/contents are validated.
'✓ All checks passed — issue is fixed' printed by the repro script; Training finished: CV AUC 0.912, submission.csv written with 10000 rows; Format check: columns, id coverage, value ranges all OK
feedback
B03
Success with no pass/fail verdict
The command exits cleanly but yields only diagnostic dumps, a side-effect confirmation (file patched), silence, or truncated/empty output, so correctness must be inferred.
Printed SQL, reprs and attribute dumps for inspection, no assertion made; 'Patch applied successfully' from a file-rewriting helper; rc=0 with empty output after grep/tail piping
feedback
B04
Environment not ready or not configured
The run aborts before exercising any logic because a required package/tool is not installed or a framework was never bootstrapped (settings, env var, app registry, DB fixtures).
ModuleNotFoundError: No module named 'pytest'; ImproperlyConfigured: settings are not configured before use; AppRegistryNotReady / missing DJANGO_SETTINGS_MODULE
feedback, rubric
B05
Unresolved name, import path, or test selector
A specific target cannot be resolved: an imported symbol or attribute does not exist, a name is used without being defined/imported, the test label/path is unknown, or the CLI invocation itself is malformed.
ImportError: cannot import name 'foo' from 'pkg.mod'; unittest.loader._FailedTest: module has no attribute 'TestX'; NameError: name 'helper' is not defined / unrecognized flag, usage error
feedback, rubric
B06
Source fails to parse, compile, or encode
Execution never happens (or dies while printing) because a file is syntactically invalid, indentation is broken by an edit, or output cannot be encoded by the console codec.
SyntaxError: unterminated string literal in the heredoc script; IndentationError after the patch reparented a block; UnicodeEncodeError while printing '✓' to an ASCII stdout
feedback, rubric
B07
Timeout, hang, or resource exhaustion
The process produces no usable verdict because it exceeded the wall-clock limit, hung on an interactive prompt or infinite loop, or was killed for running out of memory/disk/shared memory.
Command killed after 600s (rc=124) during model training; Process terminated with rc=137 'Killed' — out of memory; pip install exceeded the time limit; run dropped into pdb and never returned
feedback
B08
Missing, corrupt, or mismatched input data
The run fails on the data or files it consumes: a path does not exist, a file cannot be parsed, a referenced column/key is absent, or array shapes and lengths do not line up.
FileNotFoundError: submission.csv not found; ParserError: tokenizing data / cannot identify image file; KeyError: 'target' — column missing; shapes (1000,) and (998,) misaligned
feedback, rubric
B09
API misuse or violated library precondition
A call is rejected by the callee's contract: unexpected/missing keyword, wrong arity or argument order, an unsupported option value, or data that violates a documented precondition of the estimator/helper.
TypeError: __init__() got an unexpected keyword argument 'copy_x'; InvalidParameterError: 'auto' is no longer a supported value; ValueError: class labels must be contiguous / n_splits greater than members in a class
feedback, rubric
B10
Uncaught runtime exception on a reachable path
Exercised code raises an unhandled exception from an unguarded attribute, key, index, type, or None value, whether observed as a traceback or spotted as a guaranteed crash in the patch.
AttributeError: 'NoneType' object has no attribute 'name' inside the library frame; IndexError from a hard-coded subscript into a variable-length list; TypeError: unsupported operand type(s) raised deep in the patched function
feedback, rubric
B11
Assertion or test failure (expected vs actual)
Tests or scripts execute and report a concrete mismatch: a failing assertEqual/assert, a wrong exception or warning type, an image/golden comparison beyond tolerance, or a FAILED summary count.
Nothing raises, but the produced value is incorrect by construction or visibly wrong — inverted predicate, off-by-one/ordering slip, wrong operand or key, or a degenerate/chance-level output.
Script exits 0 but prints 'expected 3, got 4 — FAIL'; Off-by-one boundary: interior cases right, edge index past the end; All predictions collapse to a single constant / accuracy at chance level
feedback, rubric
B13
Fix misses the defect (wrong site, no-op)
The patch edits a sibling routine, unreached branch, comment/annotation, or a semantically equivalent expression — or nothing shippable at all — so the reported behavior is unchanged.
Only comments, docstrings and formatting changed; the faulty constant is untouched; Guard added in a caller while the function the report blames is byte-identical; Change exists only in a scratch script or .patch file, never applied to the source
rubric
B14
Incomplete fix: sibling paths or dead option
Only one of several parallel branches, call sites, keys, or duplicate definitions is repaired, or a newly accepted parameter/flag is never read, forwarded, or admitted by its gate.
One dispatch arm fixed, the mirrored arm still reproduces the bug; Signature gains `force=True` but the body never reads it; Fix special-cases the reported value while the rest of the named class stays broken
rubric
B15
Over-broad change or broken contract
The change reaches beyond the reported case — widened guards, altered defaults or output format, rewritten shared helpers, removed/renamed public names — regressing callers nobody complained about.
Per-case option turned into a global default, changing already-correct inputs; Public keyword argument deleted, so existing callers raise TypeError; Adjacent working call sites rewritten, breaking behavior pinned by existing tests
rubric
B16
Symptom masked or failure swallowed
The complaint is silenced rather than fixed: broad except clauses, defensive clamps or fallbacks at the crash site, discarded exit codes, or deleted/weakened guards, validators, and tests.
try/except around the failing call that only prints the error; hasattr/clamp guard added at the consumer while the mis-sized producer is untouched; Failing test deleted or its assertions stripped so the suite goes green
rubric
B17
Verification that cannot fail or misses target
The submitted check proves nothing: it prints instead of asserting, restates current behavior, uses stand-ins or the wrong copy of the code, never constructs the triggering condition, or dies in its own scaffolding.
Script prints 'SUCCESS' unconditionally with no assert or non-zero exit; Test asserts against a locally re-implemented copy of the function; Repro builds only default inputs, so it passes identically on unfixed code
rubric
B18
Repository and test-suite hygiene violations
The change set leaves scratch scripts, generated artifacts, or backup copies in the tree, drops files matching the test-discovery glob that mutate global state at import, or breaks lint/formatting conventions.
test_check.py at repo root that calls sys.exit and os.chdir at import time; One-off source-patching script and .orig backup committed with the fix; Trailing newline removed in a lint-gated repo; unused imports left behind
rubric
B19
Deliverable or answer-format mismatch
The required artifact is missing, written to the wrong path, or its schema/row alignment/answer string deviates from the supplied template or literal answer format.
submission.csv has renamed columns and a saved index, unlike sample_submission.csv; Predictions cover 900 of 1000 test ids, in a different order; Final answer hand-typed with extra quotes/brackets instead of the prescribed token
rubric
B20
Unsound analysis or validation methodology
The result is not trustworthy because the prescribed spec/data was ignored or fabricated, rows or scope were silently changed, or quality is claimed without held-out validation, baselines, or sanity checks.
Model fit on all rows and shipped with only an in-sample score; Referenced README/config never opened; bins and labels hard-coded from guesswork; dropna silently removes half the rows before the reported statistic is computed
rubric
What this says.SWE-bench Verified: mismatch 73/100 against error-like feedback (85 against all feedback). Errors the agent meets far more often than the library addresses: Success with no pass/fail verdict (6% vs 0%), Verified successful run or artifact (6% vs 0%), Assertion or test failure (expected vs actual) (25% vs 3%), Environment not ready or not configured (18% vs 0%), Unresolved name, import path, or test selector (12% vs 2%). Rubric mass without a matching test-time signal: Silently wrong value or degenerate result (15% vs 1%), Incomplete fix: sibling paths or dead option (7% vs 0%), Fix misses the defect (wrong site, no-op) (21% vs 0%), Symptom masked or failure swallowed (6% vs 0%), Verification that cannot fail or misses target (15% vs 0%). MLE-bench Lite: mismatch 98/100 against error-like feedback (97 against all feedback). Errors the agent meets far more often than the library addresses: Timeout, hang, or resource exhaustion (57% vs 0%), Uncaught runtime exception on a reachable path (18% vs 0%), Environment not ready or not configured (9% vs 0%). Rubric mass without a matching test-time signal: Silently wrong value or degenerate result (15% vs 2%), Deliverable or answer-format mismatch (34% vs 0%), Unsound analysis or validation methodology (44% vs 0%).
Test-time feedback vs rubric coverage (data-derived taxonomy)
How the buckets were made. Bottom-up from the corpus, not by hand. Corpus = every observation the interpreter agent received from a command the router classes as code execution (SWE-bench interp-c: 12,710 observations; MLE-bench mle-interp-c + mle-interp: 1,238; successes included) plus every rubric in the libraries the hybrid arms retrieve from (mined 1000 for SWE, on-policy all for MLE). Opus open-coded 48 stratified chunks of 50 items (feedback and rubric chunks, both benchmarks) into candidate categories, then merged all candidates into exactly 20 mutually exclusive buckets with definitions and examples. Haiku then labelled every observation and every rubric with one bucket (rubrics: plus an optional secondary). Code: analysis/taxonomy_induce.py.
Reading the charts. Three bars per bucket: all interpreter feedback (what the agent hears back, successes included), the error-like subset (non-zero return code or an error signature), and the rubrics (by the signal a violation would produce). Mismatch = total variation distance between two distributions in percentage points (0 = identical, 100 = disjoint).
SWE-bench Verified: share per bucket
12,710 feedback observations (3,482 error-like) from 500 instances; 1,000 rubrics. Mismatch rubrics vs all feedback 87, vs error-like feedback 79.
MLE-bench Lite: share per bucket
1,238 feedback observations (397 error-like) from 44 instances; 1,076 rubrics. Mismatch rubrics vs all feedback 99, vs error-like feedback 100.
The 20 induced buckets
id
bucket
definition
examples
sides
B01
Harness failure, timeout, or no usable output
The command produced no interpretable result because the execution infrastructure errored, the process was killed by the time limit, or the captured output is empty/truncated with no verdict.
rc=-1, 'Error executing command in GKE pod' — the command never really ran; Killed after the wall-clock limit (rc=124) mid-training, no final result; Non-zero exit with empty stdout/stderr, or only teardown noise in the captured tail
feedback
B02
Missing module, file, or unresolvable target
Execution aborts before any project logic runs because an imported package/symbol, a native library, an invoked file, or the requested test selector cannot be resolved.
ModuleNotFoundError: No module named 'pytest' / settings module not importable; ImportError: cannot import name X from Y; libGL.so.1 missing for a native extension; Runner reports unknown test id / 'can't open file' / zero tests collected / unrecognized flag
feedback
B03
Environment, config, or data prerequisite not ready
Imports succeed but the run aborts because framework configuration, database/schema state, fixtures, or an expected input/intermediate file is missing or invalid.
ImproperlyConfigured: settings are not configured / model not in INSTALLED_APPS; no such table / IntegrityError while setting up test data; FileNotFoundError for the dataset, checkpoint, or prior submission file
feedback
B04
Exception raised inside the code under test
The reproduction or test runs far enough to trigger an unhandled traceback originating in the project's or library's own code, exposing the bug or an incomplete fix.
Traceback ends inside django/db/models/... with a FieldError; Repro script raises TypeError from the library function being patched; Test ends in ERROR (exception escaped) rather than a failed assertion
feedback
B05
Scratch script dies in its own scaffolding
The agent's throwaway repro/verification script fails for reasons of its own making — guessed constructor or private API, undefined name, bad quoting, encoding of printed output — so the intended check never executes.
NameError in the heredoc script; UnicodeEncodeError while printing ✓ characters; Script crashes building its fixture with a guessed API signature before reaching the call under test; Reproduction aborts in setup scaffolding, so the target code path is never reached
feedback, rubric
B06
Type, shape, or API contract mismatch
A call boundary is violated: wrong keyword/arity for the installed API, missing column/key, wrong dtype or array shape, unseen category, or a value that may be None/short used without a guard.
TypeError: __init__() got an unexpected keyword argument 'n_estimators'; KeyError on a column the loaded dataframe does not contain; shapes (n,) vs (n,m) mismatch; Patch assumes an attribute/return shape the library does not provide, so the first real use raises
feedback, rubric
B07
Assertion or check failed (expected vs actual)
Code ran to completion but a test or hand-written check compared values and mismatched, including 'DID NOT RAISE', baseline-image diffs, and runs mixing passes with one or more failures.
FAILED test_x - AssertionError: assert 3 == 4; DID NOT RAISE <class 'ValueError'>; Script prints 4 checks PASS and 1 FAIL / suite summary '2 failed, 30 passed'
feedback
B08
All targeted tests passed
A real test-runner invocation completed with every selected test passing (possibly with skips, xfails, or warnings) and no failures or errors.
OK (skipped=2) / '35 passed, 1 warning in 4.2s'; Ran 12 tests ... OK; Selected regression tests all green after the patch
feedback
B09
Script ran and reported success or results
An ad-hoc verification script, pipeline, or validation run completes cleanly and reports the expected behavior, metrics, or a written output artifact.
'All tests passed! ✓' from the hand-written scenario script; Training finished, CV AUC 0.87, submission.csv written with 5000 rows; Format check confirms columns, id set, value range and no NaNs
feedback
B10
Silent success or diagnostic-only output
The command exits cleanly but gives no pass/fail verdict: it only applied an edit, installed a package, compiled, or printed state for inspection.
'File updated successfully' from a patch-applying script, nothing else; pip install finished with only root-user/upgrade notices; Prints dtypes, SQL, attribute dumps or source excerpts with no assertion
feedback
B11
Runs clean but the result is wrong
No exception is raised, yet the produced value or behavior is incorrect — printed output contradicts expectations, or the patch's logic (inverted condition, off-by-one, wrong constant/operand, mis-ordered statements, loop/state bug) silently computes the wrong answer.
Exit 0 but output shows 'expected 5, got 0' / NaNs still present in the predictions; Guard's comparison sense is flipped, so the interesting case returns the sentinel; Early return skips finalization; accumulator never populated, so an empty result is emitted
feedback, rubric
B12
Edited source cannot parse, import, or resolve names
The change leaves the file unusable at load time: syntax/indentation damage, leftover conflict markers, a missing import, a deleted public symbol, or a name read on a path where it is never bound.
SyntaxError/IndentationError when compiling the file the agent just rewrote; Module references an alias or helper that is never imported -> NameError at import; UnboundLocalError: assignment lives in a sibling branch or after the use
rubric, feedback
B13
No-op patch or fix at the wrong site
The submission cannot change the reported behavior: cosmetic/equivalent rewrite, edits to a function or layer the reproducer never reaches, a new flag or parameter never read, or no source change at all.
Only comments, formatting, or an equivalent expression changed; the faulty construct is byte-identical; Patch edits a neighbouring helper while the routine named in the traceback is untouched; New keyword argument accepted in the signature but never consulted in the body
rubric
B14
Symptom suppressed instead of root cause fixed
The failure is silenced rather than corrected — defensive guards or broadened except clauses at the consumer, errors swallowed and placeholders returned, or tests/harness edited or deleted to make the run green.
hasattr/try-except wrapper hides the crash while the bad value is still produced upstream; Coercion failure absorbed, field dropped, caller sees a success-shaped wrong result; Pre-existing test module emptied or sleeps added so the suite passes
rubric
B15
Incomplete fix: sibling paths left broken
Only the exact reproducer, one branch, one overload, or one of several co-reported symptoms is repaired while structurally identical call sites, subclasses, or near-miss inputs keep the old behavior.
One dunder/override fixed, the analogous peers still reproduce the defect; Per-symptom edits that skip the shared helper both symptoms flow through; Second defect named in the issue never addressed
rubric
B16
Over-broad change or collateral regression
The diff alters behavior beyond the reported scope — removing validation, options, or documented rules, widening shared defaults, changing messages/formats, or leaving out-of-diff callers, docs, and mirrored declarations inconsistent.
Fix applied in the shared unconditional path, changing results for callers that were already correct; Public keyword parameter or precedence rule deleted with no replacement; Signature narrowed without updating other call sites, __all__, or the docstring that states the old value
rubric
B17
Verification cannot fail or tests the wrong thing
The evidence offered for correctness is non-discriminating: prints instead of assertions, exceptions swallowed, expectations rewritten to current output, or the check exercises a different copy, entry point, or input than the reported case.
Script prints 'ALL CHECKS PASSED' unconditionally; assert wrapped in a bare except; Re-implements the logic locally or imports the installed package instead of the edited source; Asserts guessed literals, or cites the already-green pre-existing suite as proof
rubric
B18
Test and repository hygiene violations
The change set pollutes the tree or the suite: scratch files matching test discovery with import-time side effects, backups/generated artifacts committed, wall-clock sleeps, or unrestored global state.
Root-level test_repro.py collected by pytest runs chdir/sys.exit at module scope; .orig/.bak duplicates, patcher scripts, or generated reports left in the repo; Test sleeps on real time or mutates a process-wide singleton without teardown
rubric
B19
Deliverable or format spec not followed
The required artifact is missing, written to an ad-hoc path, or its schema/row alignment/answer string deviates from the provided template or spec file, which was often never read.
submission.csv has renamed/extra columns, an index column, or a row count that does not match the test set; Result reported only in prose; no file written at the expected path; README/sample-output that defines bins, labels, or the answer template was never opened
rubric
B20
Unsound data handling or model validation
The analysis or ML pipeline is methodologically invalid: fabricated or substituted inputs, wrong population/scope/unit of analysis, silent row loss, no held-out validation or baseline, off-metric selection, leakage, or a degenerate result accepted uncritically.
Predictions shipped with only in-sample scores; no split, CV, or trivial baseline; Hard-coded/synthetic data used because the real file was not found; required filter or grouping ignored; Cluster count chosen by bare argmax yields singleton clusters, reported as success
rubric
What this says.SWE-bench Verified: mismatch 79/100 against error-like feedback (87 against all feedback). Errors the agent meets far more often than the library addresses: Script ran and reported success or results (5% vs 0%), Missing module, file, or unresolvable target (17% vs 1%), Assertion or check failed (expected vs actual) (20% vs 3%), Harness failure, timeout, or no usable output (16% vs 0%), Scratch script dies in its own scaffolding (12% vs 4%), Environment, config, or data prerequisite not ready (10% vs 0%). Rubric mass without a matching test-time signal: Runs clean but the result is wrong (13% vs 0%), Incomplete fix: sibling paths left broken (7% vs 0%), Symptom suppressed instead of root cause fixed (7% vs 0%), No-op patch or fix at the wrong site (22% vs 0%), Over-broad change or collateral regression (6% vs 0%), Verification cannot fail or tests the wrong thing (15% vs 0%). MLE-bench Lite: mismatch 100/100 against error-like feedback (99 against all feedback). Errors the agent meets far more often than the library addresses: Harness failure, timeout, or no usable output (63% vs 0%), Exception raised inside the code under test (15% vs 0%), Missing module, file, or unresolvable target (9% vs 0%). Rubric mass without a matching test-time signal: Deliverable or format spec not followed (37% vs 0%), Unsound data handling or model validation (53% vs 0%).