Hybrid CWM: efficiency on solved instances, and what the rubrics cover

SWE-bench Verified and MLE-bench Lite · efficiency measured only where the method actually solved the task · test-time errors vs rubric coverage in one 20-bucket taxonomy · built 2026-09-16 09:21

Why filter. Unfiltered efficiency numbers reward arms that give up: an agent that submits a wrong patch after five commands has a short wall clock and a low bill. Every efficiency statistic on this page is therefore reported on the instances the arm solved, and, for the fairest reading, on the instances both the arm and the interpreter solved, so the same problems are being compared.

SWE-bench Verified (500), Haiku 4.5 agent

Setup. Per instance, wall clock = first agent command to the end of the last one, from timing.json. Sandbox time = summed duration of commands that reached the sandbox (in hybrid arms a routed command only contributes its git diff). Execution / judge time = real execution time plus the judge's latency; judge latency is measured from the timing records where the harness recorded it (harness 2 arms) and estimated otherwise (see the note under that chart). Agent time = the rest (LLM calls and harness). Calls and $ from the run logs; judge $ = verdict count × that arm's logged $ per grade.
Filters. all: every instance the arm ran. solved by the arm: the official FAIL_TO_PASS / PASS_TO_PASS grade. solved by both: the arm and the interpreter both solved it, so the same problems are compared. An arm that breaks early on hard problems looks fast on 'all' and loses that advantage under the paired filter.

Median wall clock per instance (minutes) under the three filters

all instancessolved by the armsolved by both the arm and the interpreter0.03.06.09.012.015.03.02.62.6interpreter339 solved of 5003.93.53.4floor: no execution297 solved of 50011.310.09.7hybrid CWMOpus judgemined 1000313 solved of 5004.43.93.7hybrid CWMOpus judgemined 1000 + preamble315 solved of 50011.610.510.2hybrid CWMHaiku judgemined 1000307 solved of 500

Paired n: interpreter: 339, floor: no execution: 269, hybrid CWM · Opus judge · mined 1000: 281, hybrid CWM · Opus judge · mined 1000 + preamble: 280, hybrid CWM · Haiku judge · mined 1000: 277.

Wall clock per instance, solved-by-both instances (p5–p95 whiskers, p25–p75 box, median line)

0 min7 min13 min20 min26 min33 mininterpretermedian 2.6 · n=339floor: no executionmedian 3.4 · n=269hybrid CWM · Opus judge · mined 1000median 9.7 · n=281hybrid CWM · Opus judge · mined 1000 + preamblemedian 3.7 · n=280hybrid CWM · Haiku judge · mined 1000median 10.2 · n=277

Time on code-execution commands, solved-by-both: real execution (interpreter) or estimated judge latency (hybrid)

0 min3 min6 min9 min12 min15 mininterpretermedian 0.5 · n=339floor: no executionmedian 0.0 · n=269hybrid CWM · Opus judge · mined 1000median 3.2 · n=281hybrid CWM · Opus judge · mined 1000 + preamblemedian 0.4 · n=280hybrid CWM · Haiku judge · mined 1000median 3.7 · n=277

Judge latency is not logged; it is estimated per instance as wall clock minus sandbox time minus the agent's calls x the interpreter's per-call latency (1.9 s per call on its solved instances). Arms ran on different days; Opus overload periods inflate judge latency.

Agent time (LLM calls + harness, judge estimate removed), solved-by-both instances

0 min7 min13 min20 min26 min33 mininterpretermedian 1.6 · n=339floor: no executionmedian 2.8 · n=269hybrid CWM · Opus judge · mined 1000median 2.1 · n=281hybrid CWM · Opus judge · mined 1000 + preamblemedian 2.1 · n=280hybrid CWM · Haiku judge · mined 1000median 2.2 · n=277

Agent LLM calls per instance, solved-by-both

median agent LLM calls02448729612051interpreter339 solved of 50087floor: no execution297 solved of 50067hybrid CWMOpus judgemined 1000313 solved of 50068hybrid CWMOpus judgemined 1000 + preamble315 solved of 50068hybrid CWMHaiku judgemined 1000307 solved of 500

Cost per instance, solved-by-both

median $ per instance (agent + judge)$0.00$1.00$2.00$3.00$4.00$5.00$0.19interpreter339 solved of 500$0.31floor: no execution297 solved of 500$4.01hybrid CWMOpus judgemined 1000313 solved of 500$3.62hybrid CWMOpus judgemined 1000 + preamble315 solved of 500$0.56hybrid CWMHaiku judgemined 1000307 solved of 500
armsolvedpaired nwall, allwall, solvedwall, paired{"exec + judge, paired" if judge_est else "sandbox time, paired"}judge, pairedagent time, pairedcalls, paired$, pairedrouted cmdsverdictss per routed cmd
interpreter339/5003393.02.62.60.501.651$0.191901.5
floor: no execution297/5002693.93.53.40.002.887$0.312800.0
hybrid CWM · Opus judge · mined 1000313/50028111.310.09.73.23.2 (est.)2.167$4.01191910.1
hybrid CWM · Opus judge · mined 1000 + preamble315/5002804.43.93.70.40.4 (est.)2.168$3.6217171.4
hybrid CWM · Haiku judge · mined 1000307/50027711.610.510.23.73.7 (est.)2.268$0.56202011.0

Minutes are medians. 'routed cmds' = median number of code-execution commands per instance (router class); 'verdicts' = world-model replies; 's per routed cmd' = execution (interpreter) or estimated judge latency (hybrid) per such command. Arms ran on different days, so judge latency also reflects the API load at the time (e.g. Opus overload periods); compare cost and call counts, which do not depend on that.

What this says. floor: no execution: wall 3.4 vs interpreter 2.6 min on the same 269 solved instances (+34%); exec+judge 0.0 (judge 0.0) vs execution 0.5 min; calls 87 vs 51; cost $0.31 vs $0.19
hybrid CWM · Opus judge · mined 1000: wall 9.7 vs interpreter 2.6 min on the same 281 solved instances (+278%); exec+judge 3.2 (judge 3.2 est.) vs execution 0.5 min; calls 67 vs 51; cost $4.01 vs $0.19
hybrid CWM · Opus judge · mined 1000 + preamble: wall 3.7 vs interpreter 2.6 min on the same 280 solved instances (+46%); exec+judge 0.4 (judge 0.4 est.) vs execution 0.5 min; calls 68 vs 51; cost $3.62 vs $0.19
hybrid CWM · Haiku judge · mined 1000: wall 10.2 vs interpreter 2.6 min on the same 277 solved instances (+300%); exec+judge 3.7 (judge 3.7 est.) vs execution 0.5 min; calls 68 vs 51; cost $0.56 vs $0.19

MLE-bench Lite (22), Haiku 4.5 agent, solved = beats the no-execution floor

Setup. Per instance, wall clock = first agent command to the end of the last one, from timing.json (MLE data-setup commands excluded). Sandbox time = summed duration of commands that reached the sandbox (in hybrid arms a routed command only contributes its git diff). Execution / judge time = real execution time plus the judge's latency; judge latency is measured from the timing records where the harness recorded it (harness 2 arms) and estimated otherwise (see the note under that chart). Agent time = the rest (LLM calls and harness). Calls and $ from the run logs; judge $ = verdict count × that arm's logged $ per grade.
Filters. all: every instance the arm ran. solved by the arm: a valid submission from a run that finished (exit status 'submitted') whose score, relative to the leaderboard median, is better than the no-execution floor's on the same competition by at least 0.02. This excludes constant-prediction files (valid, but identical to doing nothing) without demanding a medal. The threshold sweep below shows how every count depends on the bar. solved by both: the arm and the interpreter both solved it, so the same problems are compared. An arm that breaks early on hard problems looks fast on 'all' and loses that advantage under the paired filter.

Median wall clock per instance (minutes) under the three filters

all instancessolved by the armsolved by both the arm and the interpreter0.024.048.072.096.0120.046.542.342.3interpreter11 solved of 228.6n/an/afloor: no execution0 solved of 229.47.37.1hybrid CWMOpus judgeall rubrics7 solved of 229.79.19.8hybrid CWMOpus judge1000 rubrics9 solved of 229.67.06.7hybrid CWMHaiku judgeall rubrics6 solved of 21

Paired n: interpreter: 11, floor: no execution: 0, hybrid CWM · Opus judge · all rubrics: 5, hybrid CWM · Opus judge · 1000 rubrics: 6, hybrid CWM · Haiku judge · all rubrics: 4.

Wall clock per instance, solved-by-both instances (p5–p95 whiskers, p25–p75 box, median line)

0 min53 min106 min158 min211 min264 mininterpretermedian 42.3 · n=11floor: no executionn/ahybrid CWM · Opus judge · all rubricsmedian 7.1 · n=5hybrid CWM · Opus judge · 1000 rubricsmedian 9.8 · n=6hybrid CWM · Haiku judge · all rubricsmedian 6.7 · n=4

Sandbox time (recorded command durations), solved-by-both instances

0 min24 min48 min72 min96 min120 mininterpretermedian 38.9 · n=11floor: no executionn/ahybrid CWM · Opus judge · all rubricsmedian 2.3 · n=5hybrid CWM · Opus judge · 1000 rubricsmedian 5.4 · n=6hybrid CWM · Haiku judge · all rubricsmedian 2.7 · n=4

On MLE-bench the interpreter's remaining time per call is dominated by long data-loading and training waits, so the judge-latency estimate used on SWE-bench is not meaningful here; only recorded sandbox time is shown.

Agent LLM calls per instance, solved-by-both

median agent LLM calls02448729612042interpreter11 solved of 22n/afloor: no execution0 solved of 22114hybrid CWMOpus judgeall rubrics7 solved of 22114hybrid CWMOpus judge1000 rubrics9 solved of 22108hybrid CWMHaiku judgeall rubrics6 solved of 21

Cost per instance, solved-by-both

median $ per instance (agent + judge)$0.00$2.00$4.00$6.00$8.00$10.00$0.37interpreter11 solved of 22n/afloor: no execution0 solved of 22$8.10hybrid CWMOpus judgeall rubrics7 solved of 22$8.21hybrid CWMOpus judge1000 rubrics9 solved of 22$0.95hybrid CWMHaiku judgeall rubrics6 solved of 21
armsolvedpaired nwall, allwall, solvedwall, paired{"exec + judge, paired" if judge_est else "sandbox time, paired"}judge, pairedagent time, pairedcalls, paired$, pairedrouted cmdsverdictss per routed cmd
interpreter11/221146.542.342.338.90n/a42$0.37230100.5
floor: no execution0/2208.6n/an/an/a0n/an/an/an/an/an/a
hybrid CWM · Opus judge · all rubrics7/2259.47.37.12.30n/a114$8.1045450.0
hybrid CWM · Opus judge · 1000 rubrics9/2269.79.19.85.40n/a114$8.2134340.0
hybrid CWM · Haiku judge · all rubrics6/2149.67.06.72.70n/a108$0.9530300.0

Minutes are medians. 'routed cmds' = median number of code-execution commands per instance (router class); 'verdicts' = world-model replies; 's per routed cmd' = execution (interpreter) or estimated judge latency (hybrid) per such command. Arms ran on different days, so judge latency also reflects the API load at the time (e.g. Opus overload periods); compare cost and call counts, which do not depend on that.

Per arm: beats floor / >= 90% of median / above median / medal / valid: interpreter: 11 / 7 / 2 / 1 / 18, floor: no execution: 0 / 1 / 0 / 0 / 19, hybrid CWM · Opus judge · all rubrics: 7 / 3 / 0 / 0 / 19, hybrid CWM · Opus judge · 1000 rubrics: 9 / 2 / 0 / 0 / 19, hybrid CWM · Haiku judge · all rubrics: 6 / 0 / 0 / 0 / 18.

MLE-bench: every valid submission, score relative to the leaderboard median vs wall clock

0.000.260.520.781.041.300 min22 min43 min65 min87 min108 min130 minleaderboard mediancompetitive bar (90% of median)interpreteraerial-cactus-identification: 1.00 of median, 50 minsiim-isic-melanoma-classification: 0.56 of median, 34 minhistopathologic-cancer-detection: 0.89 of median, 64 minnomad2018-predict-transparent-conductors: 1.11 of median, 25 minranzcr-clip-catheter-line-classification: 0.53 of median, 52 minleaf-classification: 0.15 of median, 42 minthe-icml-2013-whale-challenge-right-whale-redux: 0.69 of median, 120 mintabular-playground-series-dec-2021: 1.00 of median, 91 minnew-york-city-taxi-fare-prediction: 0.79 of median, 62 minmlsp-2013-birds: 0.95 of median, 28 minjigsaw-toxic-comment-classification-challenge: 0.98 of median, 41 minplant-pathology-2020-fgvc7: 0.53 of median, 41 mindog-breed-identification: 0.10 of median, 53 mintabular-playground-series-may-2022: 0.93 of median, 49 minspooky-author-identification: 0.77 of median, 30 minrandom-acts-of-pizza: 1.08 of median, 9 minaptos2019-blindness-detection: 0.62 of median, 51 mindenoising-dirty-documents: 0.49 of median, 57 minfloor: no executionaerial-cactus-identification: 0.50 of median, 8 minsiim-isic-melanoma-classification: 0.74 of median, 8 minhistopathologic-cancer-detection: 0.53 of median, 9 minnomad2018-predict-transparent-conductors: 0.60 of median, 12 minranzcr-clip-catheter-line-classification: 0.52 of median, 8 mindetecting-insults-in-social-commentary: 0.66 of median, 12 minleaf-classification: 0.02 of median, 8 minthe-icml-2013-whale-challenge-right-whale-redux: 0.65 of median, 11 mindogs-vs-cats-redux-kernels-edition: 0.14 of median, 39 mintabular-playground-series-dec-2021: 0.63 of median, 11 minnew-york-city-taxi-fare-prediction: 0.32 of median, 8 minmlsp-2013-birds: 0.79 of median, 8 minjigsaw-toxic-comment-classification-challenge: 0.51 of median, 22 minplant-pathology-2020-fgvc7: 0.53 of median, 7 mindog-breed-identification: 0.10 of median, 8 mintabular-playground-series-may-2022: 0.94 of median, 16 minspooky-author-identification: 0.39 of median, 9 minrandom-acts-of-pizza: 0.83 of median, 8 minaptos2019-blindness-detection: 0.02 of median, 7 minhybrid CWM · Opus judge · all rubricsaerial-cactus-identification: 0.98 of median, 10 minsiim-isic-melanoma-classification: 0.73 of median, 8 minhistopathologic-cancer-detection: 0.51 of median, 63 minnomad2018-predict-transparent-conductors: 0.74 of median, 7 minranzcr-clip-catheter-line-classification: 0.52 of median, 14 mindetecting-insults-in-social-commentary: 0.91 of median, 11 minleaf-classification: 0.02 of median, 10 minthe-icml-2013-whale-challenge-right-whale-redux: 0.58 of median, 9 mintabular-playground-series-dec-2021: 0.45 of median, 5 minnew-york-city-taxi-fare-prediction: 0.41 of median, 4 minmlsp-2013-birds: 0.81 of median, 7 minjigsaw-toxic-comment-classification-challenge: 0.53 of median, 35 minplant-pathology-2020-fgvc7: 0.55 of median, 11 mindog-breed-identification: 0.10 of median, 22 mintabular-playground-series-may-2022: 0.55 of median, 4 mintext-normalization-challenge-russian-language: 0.99 of median, 5 minspooky-author-identification: 0.39 of median, 11 minrandom-acts-of-pizza: 0.86 of median, 7 minaptos2019-blindness-detection: -0.04 of median, 9 minhybrid CWM · Opus judge · 1000 rubricsaerial-cactus-identification: 0.80 of median, 19 minsiim-isic-melanoma-classification: 0.81 of median, 9 minhistopathologic-cancer-detection: 0.53 of median, 11 minnomad2018-predict-transparent-conductors: 0.59 of median, 7 minranzcr-clip-catheter-line-classification: 0.52 of median, 7 mindetecting-insults-in-social-commentary: 0.95 of median, 7 minleaf-classification: 0.02 of median, 7 minthe-icml-2013-whale-challenge-right-whale-redux: 0.58 of median, 40 mindogs-vs-cats-redux-kernels-edition: 0.18 of median, 18 mintabular-playground-series-dec-2021: 0.89 of median, 21 minnew-york-city-taxi-fare-prediction: 0.60 of median, 7 minmlsp-2013-birds: 0.77 of median, 13 minjigsaw-toxic-comment-classification-challenge: 0.83 of median, 13 minplant-pathology-2020-fgvc7: 0.53 of median, 5 mindog-breed-identification: 0.10 of median, 10 mintabular-playground-series-may-2022: 0.59 of median, 16 minspooky-author-identification: 0.38 of median, 8 minrandom-acts-of-pizza: 0.91 of median, 5 minaptos2019-blindness-detection: 0.52 of median, 5 minhybrid CWM · Haiku judge · all rubricsaerial-cactus-identification: 0.41 of median, 57 minhistopathologic-cancer-detection: 0.53 of median, 10 minnomad2018-predict-transparent-conductors: 0.54 of median, 5 minranzcr-clip-catheter-line-classification: 0.52 of median, 5 minleaf-classification: 0.02 of median, 8 minthe-icml-2013-whale-challenge-right-whale-redux: 0.60 of median, 27 mindogs-vs-cats-redux-kernels-edition: 0.18 of median, 10 mintabular-playground-series-dec-2021: 0.56 of median, 16 minnew-york-city-taxi-fare-prediction: 0.52 of median, 8 minmlsp-2013-birds: 0.85 of median, 5 minjigsaw-toxic-comment-classification-challenge: 0.51 of median, 9 minplant-pathology-2020-fgvc7: 0.57 of median, 6 mindog-breed-identification: 0.10 of median, 19 mintabular-playground-series-may-2022: 0.56 of median, 13 minspooky-author-identification: 0.39 of median, 8 minrandom-acts-of-pizza: 0.88 of median, 12 minaptos2019-blindness-detection: 0.06 of median, 6 mindenoising-dirty-documents: 0.48 of median, 26 miny: score / leaderboard median (direction-aware) · x: wall clock

One point per competition and arm (hover for the name). Points on the same horizontal line across arms are constant-prediction submissions that score like the floor. The hybrid's speed comes from never training a model; only points above the competitive bar count as solved.

Threshold sweep: competitions counted as solved at each bar (finished, valid, score/median >= bar)

interpreterfloor: no executionhybrid CWM · Opus judge · all rubricshybrid CWM · Opus judge · 1000 rubricshybrid CWM · Haiku judge · all rubrics049131822141313151150% of median11878360% of median10477270% of median10357275% of median8256280% of median8143185% of median7132090% of median20000at the median

At 50% of the median even the floor passes most competitions; at the median the hybrids pass none.

barinterpreter solvedfloor: no execution: paired n · wall hybrid vs interpreterhybrid CWM · Opus judge · all rubrics: paired n · wall hybrid vs interpreterhybrid CWM · Opus judge · 1000 rubrics: paired n · wall hybrid vs interpreterhybrid CWM · Haiku judge · all rubrics: paired n · wall hybrid vs interpreter
50% of median1411 · 8 vs 41 min10 · 9 vs 41 min13 · 9 vs 49 min10 · 8 vs 45 min
60% of median115 · 11 vs 28 min4 · 7 vs 27 min6 · 13 vs 46 min2 · 9 vs 19 min
70% of median103 · 8 vs 28 min4 · 7 vs 27 min5 · 13 vs 41 min2 · 9 vs 19 min
75% of median103 · 8 vs 28 min3 · 7 vs 28 min5 · 13 vs 41 min2 · 9 vs 19 min
80% of median82 · 12 vs 29 min3 · 7 vs 28 min4 · 16 vs 46 min2 · 9 vs 19 min
85% of median81 · 16 vs 49 min2 · 8 vs 30 min2 · 13 vs 50 min1 · 12 vs 9 min
90% of median71 · 16 vs 49 min1 · 10 vs 50 min1 · 5 vs 9 min0
at the median20000

Paired n = competitions both the arm and the interpreter pass at that bar; minutes are medians over those competitions.

What this says. Paired sets are tiny here (floor: no execution: 0, hybrid CWM · Haiku judge · all rubrics: 4 competitions solved by both), so the lines below are illustrative, not a measurement.
hybrid CWM · Opus judge · all rubrics: wall 7.1 vs interpreter 42.3 min on the same 5 solved instances (-83%); sandbox time 2.3 vs 38.9 min; calls 114 vs 42; cost $8.10 vs $0.37
hybrid CWM · Opus judge · 1000 rubrics: wall 9.8 vs interpreter 42.3 min on the same 6 solved instances (-77%); sandbox time 5.4 vs 38.9 min; calls 114 vs 42; cost $8.21 vs $0.37
hybrid CWM · Haiku judge · all rubrics: wall 6.7 vs interpreter 42.3 min on the same 4 solved instances (-84%); sandbox time 2.7 vs 38.9 min; calls 108 vs 42; cost $0.95 vs $0.37

Opus vs Haiku as the world model, same library, on the instances each arm solved that the interpreter also solved

benchmarksolved (Opus / Haiku)paired nmedian wall minmedian exec / judge min (est.)median callsmedian $
SWE-bench Verified313 / 307281 / 2779.7 / 10.23.2 / 3.767 / 68$4.01 / $0.56
MLE-bench Lite (beats the floor; exec/judge column = sandbox time)7 / 65 / 47.1 / 6.70.0 / 0.0114 / 108$8.10 / $0.95
What this says. On SWE-bench, Opus solves 313 vs Haiku 307; on the instances each solved that the interpreter also solved, Opus is faster per instance (10.2 vs 9.7 min) and Haiku is cheaper ($0.56 vs $4.01). Estimated judge time per instance: Opus 3.2 min vs Haiku 3.7. Opus remains the better judge on outcome; Haiku's advantage is only latency and cost, and both hybrids stay behind the interpreter on resolve.

Harness 2 (9/14): interpreter vs CWM-only vs dual channel

What is different. Harness pulled 2026-09-14: the agent is told which channel it is on, package installs never routed, 120 s grader timeout, reopening circuit breaker, and every judge call is timed, so judge time below is measured, not estimated. CWM only = routed commands never run, verdict only (mined 1000 / on-policy all + task criterion + preamble). Dual N s = the routed command runs for real and is killed at N seconds; the agent gets its output (or the partial output and a 'killed' line) and the verdict. MLE arms all write final.py, run once uncapped at collection. Plan: docs/PLAN_H2_0914.md.

SWE-bench Verified (500), harness 2

Setup. Per instance, wall clock = first agent command to the end of the last one, from timing.json. Sandbox time = summed duration of commands that reached the sandbox (in hybrid arms a routed command only contributes its git diff). Execution / judge time = real execution time plus the judge's latency; judge latency is measured from the timing records where the harness recorded it (harness 2 arms) and estimated otherwise (see the note under that chart). Agent time = the rest (LLM calls and harness). Calls and $ from the run logs; judge $ = verdict count × that arm's logged $ per grade.
Filters. all: every instance the arm ran. solved by the arm: the official FAIL_TO_PASS / PASS_TO_PASS grade. solved by both: the arm and the interpreter both solved it, so the same problems are compared. An arm that breaks early on hard problems looks fast on 'all' and loses that advantage under the paired filter.

Median wall clock per instance (minutes) under the three filters

all instancessolved by the armsolved by both the arm and the interpreter0.03.06.09.012.015.04.84.44.4interpreter347 solved of 5005.14.24.0floor: no execution325 solved of 5004.73.83.7CWM onlyOpus324 solved of 5005.44.34.3CWM onlyHaiku307 solved of 5006.55.15.0dual 30 sOpus346 solved of 5006.96.05.8dual 30 sHaiku340 solved of 5006.35.14.9dual 60 sOpus348 solved of 5006.95.95.8dual 60 sHaiku339 solved of 5006.35.35.0dual 120 sOpus352 solved of 5006.75.95.7dual 120 sHaiku342 solved of 500

Paired n: interpreter: 347, floor: no execution: 300, CWM only · Opus: 300, CWM only · Haiku: 291, dual 30 s · Opus: 319, dual 30 s · Haiku: 317, dual 60 s · Opus: 320, dual 60 s · Haiku: 315, dual 120 s · Opus: 324, dual 120 s · Haiku: 319.

Wall clock per instance, solved-by-both instances (p5–p95 whiskers, p25–p75 box, median line)

0 min7 min13 min20 min26 min33 mininterpretermedian 4.4 · n=347floor: no executionmedian 4.0 · n=300CWM only · Opusmedian 3.7 · n=300CWM only · Haikumedian 4.3 · n=291dual 30 s · Opusmedian 5.0 · n=319dual 30 s · Haikumedian 5.8 · n=317dual 60 s · Opusmedian 4.9 · n=320dual 60 s · Haikumedian 5.8 · n=315dual 120 s · Opusmedian 5.0 · n=324dual 120 s · Haikumedian 5.7 · n=319

Time on code-execution commands, solved-by-both: real execution (interpreter) or estimated judge latency (hybrid)

0 min3 min6 min9 min12 min15 mininterpretermedian 0.7 · n=347floor: no executionmedian 0.0 · n=300CWM only · Opusmedian 0.3 · n=300CWM only · Haikumedian 0.5 · n=291dual 30 s · Opusmedian 1.9 · n=319dual 30 s · Haikumedian 2.5 · n=317dual 60 s · Opusmedian 1.9 · n=320dual 60 s · Haikumedian 2.5 · n=315dual 120 s · Opusmedian 2.0 · n=324dual 120 s · Haikumedian 2.6 · n=319

Judge latency is not logged; it is estimated per instance as wall clock minus sandbox time minus the agent's calls x the interpreter's per-call latency (3.5 s per call on its solved instances). Arms ran on different days; Opus overload periods inflate judge latency.

Agent time (LLM calls + harness, judge estimate removed), solved-by-both instances

0 min7 min13 min20 min26 min33 mininterpretermedian 3.0 · n=347floor: no executionmedian 3.0 · n=300CWM only · Opusmedian 2.3 · n=300CWM only · Haikumedian 2.7 · n=291dual 30 s · Opusmedian 1.9 · n=319dual 30 s · Haikumedian 2.1 · n=317dual 60 s · Opusmedian 1.8 · n=320dual 60 s · Haikumedian 2.0 · n=315dual 120 s · Opusmedian 1.9 · n=324dual 120 s · Haikumedian 2.0 · n=319

Agent LLM calls per instance, solved-by-both

median agent LLM calls02448729612053interpreter347 solved of 50055floor: no execution325 solved of 50050CWM onlyOpus324 solved of 50053CWM onlyHaiku307 solved of 50040dual 30 sOpus346 solved of 50042dual 30 sHaiku340 solved of 50041dual 60 sOpus348 solved of 50042dual 60 sHaiku339 solved of 50040dual 120 sOpus352 solved of 50042dual 120 sHaiku342 solved of 500

Cost per instance, solved-by-both

median $ per instance (agent + judge)$0.00$1.20$2.40$3.60$4.80$6.00$0.20interpreter347 solved of 500$1.82floor: no execution325 solved of 500$1.46CWM onlyOpus324 solved of 500$0.33CWM onlyHaiku307 solved of 500$2.40dual 30 sOpus346 solved of 500$0.35dual 30 sHaiku340 solved of 500$2.54dual 60 sOpus348 solved of 500$0.34dual 60 sHaiku339 solved of 500$2.58dual 120 sOpus352 solved of 500$0.34dual 120 sHaiku342 solved of 500
armsolvedpaired nwall, allwall, solvedwall, paired{"exec + judge, paired" if judge_est else "sandbox time, paired"}judge, pairedagent time, pairedcalls, paired$, pairedrouted cmdsverdictss per routed cmd
interpreter347/5003474.84.44.40.703.053$0.202102.0
floor: no execution325/5003005.14.24.00.003.055$1.82890.0
CWM only · Opus324/5003004.73.83.70.30.3 (est.)2.350$1.46564.1
CWM only · Haiku307/5002915.44.34.30.50.5 (est.)2.753$0.33674.9
dual 30 s · Opus346/5003196.55.15.01.91.41.940$2.4013118.9
dual 30 s · Haiku340/5003176.96.05.82.52.12.142$0.35141210.9
dual 60 s · Opus348/5003206.35.14.91.91.51.841$2.5413119.0
dual 60 s · Haiku339/5003156.95.95.82.52.12.042$0.34131211.7
dual 120 s · Opus352/5003246.35.35.02.01.51.940$2.5813119.2
dual 120 s · Haiku342/5003196.75.95.72.62.02.042$0.34131211.9

Minutes are medians. 'routed cmds' = median number of code-execution commands per instance (router class); 'verdicts' = world-model replies; 's per routed cmd' = execution (interpreter) or estimated judge latency (hybrid) per such command. Arms ran on different days, so judge latency also reflects the API load at the time (e.g. Opus overload periods); compare cost and call counts, which do not depend on that.

What this says. floor: no execution: wall 4.0 vs interpreter 4.4 min on the same 300 solved instances (-10%); exec+judge 0.0 (judge 0.0 est.) vs execution 0.7 min; calls 55 vs 53; cost $1.82 vs $0.20
CWM only · Opus: wall 3.7 vs interpreter 4.4 min on the same 300 solved instances (-17%); exec+judge 0.3 (judge 0.3 est.) vs execution 0.7 min; calls 50 vs 53; cost $1.46 vs $0.20
CWM only · Haiku: wall 4.3 vs interpreter 4.4 min on the same 291 solved instances (-4%); exec+judge 0.5 (judge 0.5 est.) vs execution 0.7 min; calls 53 vs 53; cost $0.33 vs $0.20
dual 30 s · Opus: wall 5.0 vs interpreter 4.4 min on the same 319 solved instances (+13%); exec+judge 1.9 (judge 1.4) vs execution 0.7 min; calls 40 vs 53; cost $2.40 vs $0.20
dual 30 s · Haiku: wall 5.8 vs interpreter 4.4 min on the same 317 solved instances (+31%); exec+judge 2.5 (judge 2.1) vs execution 0.7 min; calls 42 vs 53; cost $0.35 vs $0.20
dual 60 s · Opus: wall 4.9 vs interpreter 4.4 min on the same 320 solved instances (+11%); exec+judge 1.9 (judge 1.5) vs execution 0.7 min; calls 41 vs 53; cost $2.54 vs $0.20
dual 60 s · Haiku: wall 5.8 vs interpreter 4.4 min on the same 315 solved instances (+31%); exec+judge 2.5 (judge 2.1) vs execution 0.7 min; calls 42 vs 53; cost $0.34 vs $0.20
dual 120 s · Opus: wall 5.0 vs interpreter 4.4 min on the same 324 solved instances (+14%); exec+judge 2.0 (judge 1.5) vs execution 0.7 min; calls 40 vs 53; cost $2.58 vs $0.20
dual 120 s · Haiku: wall 5.7 vs interpreter 4.4 min on the same 319 solved instances (+28%); exec+judge 2.6 (judge 2.0) vs execution 0.7 min; calls 42 vs 53; cost $0.34 vs $0.20

MLE-bench Lite (22), harness 2, solved = beats the harness-2 floor

Setup. Per instance, wall clock = first agent command to the end of the last one, from timing.json (MLE data-setup commands excluded). Sandbox time = summed duration of commands that reached the sandbox (in hybrid arms a routed command only contributes its git diff). Execution / judge time = real execution time plus the judge's latency; judge latency is measured from the timing records where the harness recorded it (harness 2 arms) and estimated otherwise (see the note under that chart). Agent time = the rest (LLM calls and harness). Calls and $ from the run logs; judge $ = verdict count × that arm's logged $ per grade.
Filters. all: every instance the arm ran. solved by the arm: finished, valid, score/median better than h2-mle-floor's on the same competition by 0.02. solved by both: the arm and the interpreter both solved it, so the same problems are compared. An arm that breaks early on hard problems looks fast on 'all' and loses that advantage under the paired filter.

Median wall clock per instance (minutes) under the three filters

all instancessolved by the armsolved by both the arm and the interpreter0.024.048.072.096.0120.044.868.768.7interpreter4 solved of 2220.1n/an/afloor: no execution0 solved of 2216.930.638.0CWM onlyOpus2 solved of 2230.321.521.5CWM onlyHaiku1 solved of 2229.2n/an/adual 30 sOpus0 solved of 2226.327.327.3dual 30 sHaiku1 solved of 2225.713.017.7dual 60 sOpus2 solved of 2232.150.185.8dual 60 sHaiku2 solved of 2228.0n/an/adual 120 sOpus0 solved of 2225.225.925.9dual 120 sHaiku3 solved of 22

Paired n: interpreter: 4, floor: no execution: 0, CWM only · Opus: 1, CWM only · Haiku: 1, dual 30 s · Opus: 0, dual 30 s · Haiku: 1, dual 60 s · Opus: 1, dual 60 s · Haiku: 1, dual 120 s · Opus: 0, dual 120 s · Haiku: 3.

Wall clock per instance, solved-by-both instances (p5–p95 whiskers, p25–p75 box, median line)

0 min53 min106 min158 min211 min264 mininterpretermedian 68.7 · n=4floor: no executionn/aCWM only · Opusmedian 38.0 · n=1CWM only · Haikumedian 21.5 · n=1dual 30 s · Opusn/adual 30 s · Haikumedian 27.3 · n=1dual 60 s · Opusmedian 17.7 · n=1dual 60 s · Haikumedian 85.8 · n=1dual 120 s · Opusn/adual 120 s · Haikumedian 25.9 · n=3

Sandbox time (recorded command durations), solved-by-both instances

0 min24 min48 min72 min96 min120 mininterpretermedian 60.2 · n=4floor: no executionn/aCWM only · Opusmedian 28.3 · n=1CWM only · Haikumedian 7.4 · n=1dual 30 s · Opusn/adual 30 s · Haikumedian 5.6 · n=1dual 60 s · Opusmedian 9.0 · n=1dual 60 s · Haikumedian 18.3 · n=1dual 120 s · Opusn/adual 120 s · Haikumedian 11.3 · n=3

On MLE-bench the interpreter's remaining time per call is dominated by long data-loading and training waits, so the judge-latency estimate used on SWE-bench is not meaningful here; only recorded sandbox time is shown.

Agent LLM calls per instance, solved-by-both

median agent LLM calls02448729612063interpreter4 solved of 22n/afloor: no execution0 solved of 2256CWM onlyOpus2 solved of 2268CWM onlyHaiku1 solved of 22n/adual 30 sOpus0 solved of 2291dual 30 sHaiku1 solved of 2254dual 60 sOpus2 solved of 22244dual 60 sHaiku2 solved of 22n/adual 120 sOpus0 solved of 2238dual 120 sHaiku3 solved of 22

Cost per instance, solved-by-both

median $ per instance (agent + judge)$0.00$2.00$4.00$6.00$8.00$10.00$0.58interpreter4 solved of 22n/afloor: no execution0 solved of 22$3.53CWM onlyOpus2 solved of 22$0.77CWM onlyHaiku1 solved of 22n/adual 30 sOpus0 solved of 22$1.09dual 30 sHaiku1 solved of 22$4.49dual 60 sOpus2 solved of 22$2.54dual 60 sHaiku2 solved of 22n/adual 120 sOpus0 solved of 22$0.64dual 120 sHaiku3 solved of 22
armsolvedpaired nwall, allwall, solvedwall, paired{"exec + judge, paired" if judge_est else "sandbox time, paired"}judge, pairedagent time, pairedcalls, paired$, pairedrouted cmdsverdictss per routed cmd
interpreter4/22444.868.768.760.20n/a63$0.58200182.1
floor: no execution0/22020.1n/an/an/a0n/an/an/an/an/an/a
CWM only · Opus2/22116.930.638.028.32.3n/a56$3.531313135.1
CWM only · Haiku1/22130.321.521.57.43.7n/a68$0.77141442.0
dual 30 s · Opus0/22029.2n/an/an/a0n/an/an/an/an/an/a
dual 30 s · Haiku1/22126.327.327.35.63.2n/a91$1.09161429.9
dual 60 s · Opus2/22125.713.017.79.02.3n/a54$4.49211730.8
dual 60 s · Haiku2/22132.150.185.818.34.1n/a244$2.54242049.9
dual 120 s · Opus0/22028.0n/an/an/a0n/an/an/an/an/an/a
dual 120 s · Haiku3/22325.225.925.911.34.3n/a38$0.64181652.2

Minutes are medians. 'routed cmds' = median number of code-execution commands per instance (router class); 'verdicts' = world-model replies; 's per routed cmd' = execution (interpreter) or estimated judge latency (hybrid) per such command. Arms ran on different days, so judge latency also reflects the API load at the time (e.g. Opus overload periods); compare cost and call counts, which do not depend on that.

What this says. Paired sets are tiny here (floor: no execution: 0, CWM only · Opus: 1, CWM only · Haiku: 1, dual 30 s · Opus: 0, dual 30 s · Haiku: 1, dual 60 s · Opus: 1, dual 60 s · Haiku: 1, dual 120 s · Opus: 0, dual 120 s · Haiku: 3 competitions solved by both), so the lines below are illustrative, not a measurement.
CWM only · Opus: wall 38.0 vs interpreter 68.7 min on the same 1 solved instances (-45%); sandbox time 28.3 vs 60.2 min; calls 56 vs 63; cost $3.53 vs $0.58
CWM only · Haiku: wall 21.5 vs interpreter 68.7 min on the same 1 solved instances (-69%); sandbox time 7.4 vs 60.2 min; calls 68 vs 63; cost $0.77 vs $0.58
dual 30 s · Haiku: wall 27.3 vs interpreter 68.7 min on the same 1 solved instances (-60%); sandbox time 5.6 vs 60.2 min; calls 91 vs 63; cost $1.09 vs $0.58
dual 60 s · Opus: wall 17.7 vs interpreter 68.7 min on the same 1 solved instances (-74%); sandbox time 9.0 vs 60.2 min; calls 54 vs 63; cost $4.49 vs $0.58
dual 60 s · Haiku: wall 85.8 vs interpreter 68.7 min on the same 1 solved instances (+25%); sandbox time 18.3 vs 60.2 min; calls 244 vs 63; cost $2.54 vs $0.58
dual 120 s · Haiku: wall 25.9 vs interpreter 68.7 min on the same 3 solved instances (-62%); sandbox time 11.3 vs 60.2 min; calls 38 vs 63; cost $0.64 vs $0.58

Test-time feedback vs rubric coverage, harness-2 traces (buckets re-induced from these traces)

How the buckets were made. Bottom-up from the corpus, not by hand. Corpus = every observation the interpreter agent received from a command the router classes as code execution (harness-2 interpreter traces: SWE-bench h2-interp and MLE-bench h2-mle-interp, successes included) plus every rubric in the libraries the hybrid arms retrieve from (mined 1000 for SWE, on-policy all for MLE). Opus open-coded 48 stratified chunks of 50 items (feedback and rubric chunks, both benchmarks) into candidate categories, then merged all candidates into exactly 20 mutually exclusive buckets with definitions and examples. Haiku then labelled every observation and every rubric with one bucket (rubrics: plus an optional secondary). Code: analysis/taxonomy_induce.py.
Reading the charts. Three bars per bucket: all interpreter feedback (what the agent hears back, successes included), the error-like subset (non-zero return code or an error signature), and the rubrics (by the signal a violation would produce). Mismatch = total variation distance between two distributions in percentage points (0 = identical, 100 = disjoint).

SWE-bench Verified: share per bucket

all interpreter feedbackerror-like feedback onlyrubricsSuccess with no pass/fail verdict28.3% (3637)5.8% (182)0.1% (1)Verified successful run or artifact25.4% (3262)6.3% (197)0.0% (0)All tests pass19.9% (2557)1.8% (55)0.0% (0)Assertion or test failure (expected vs actual)7.0% (895)25.0% (781)2.7% (27)Environment not ready or not configured6.0% (771)18.4% (575)0.5% (5)Uncaught runtime exception on a reachable path4.8% (617)19.1% (597)17.2% (172)Unresolved name, import path, or test selector3.3% (425)12.4% (388)2.4% (24)Silently wrong value or degenerate result1.5% (194)0.8% (24)14.8% (148)Source fails to parse, compile, or encode1.5% (193)5.9% (184)1.7% (17)Missing, corrupt, or mismatched input data1.3% (171)2.4% (76)0.5% (5)API misuse or violated library precondition0.5% (67)0.9% (27)1.8% (18)Timeout, hang, or resource exhaustion0.3% (34)1.1% (33)0.3% (3)Incomplete fix: sibling paths or dead option0.1% (8)0.2% (5)6.8% (68)Fix misses the defect (wrong site, no-op)0.0% (0)0.0% (0)21.2% (212)Over-broad change or broken contract0.0% (0)0.0% (0)4.8% (48)Symptom masked or failure swallowed0.0% (0)0.0% (0)6.2% (62)Verification that cannot fail or misses target0.0% (0)0.0% (0)14.9% (149)Repository and test-suite hygiene violations0.0% (0)0.0% (0)3.7% (37)Deliverable or answer-format mismatch0.0% (0)0.0% (0)0.1% (1)Unsound analysis or validation methodology0.0% (0)0.0% (0)0.3% (3)

12,840 feedback observations (3,126 error-like) from 500 instances; 1,000 rubrics. Mismatch rubrics vs all feedback 85, vs error-like feedback 73.

MLE-bench Lite: share per bucket

all interpreter feedbackerror-like feedback onlyrubricsVerified successful run or artifact61.6% (373)0.8% (1)0.0% (0)Success with no pass/fail verdict14.5% (88)0.0% (0)0.0% (0)Timeout, hang, or resource exhaustion12.2% (74)56.9% (74)0.0% (0)Uncaught runtime exception on a reachable path3.8% (23)17.7% (23)0.0% (0)Environment not ready or not configured2.1% (13)9.2% (12)0.0% (0)Silently wrong value or degenerate result2.1% (13)1.5% (2)15.2% (164)API misuse or violated library precondition1.0% (6)4.6% (6)0.0% (0)Source fails to parse, compile, or encode0.8% (5)1.5% (2)0.1% (1)Missing, corrupt, or mismatched input data0.8% (5)3.8% (5)0.6% (6)Assertion or test failure (expected vs actual)0.7% (4)2.3% (3)0.1% (1)Unresolved name, import path, or test selector0.3% (2)1.5% (2)0.1% (1)All tests pass0.0% (0)0.0% (0)0.0% (0)Fix misses the defect (wrong site, no-op)0.0% (0)0.0% (0)3.8% (41)Incomplete fix: sibling paths or dead option0.0% (0)0.0% (0)0.2% (2)Over-broad change or broken contract0.0% (0)0.0% (0)0.0% (0)Symptom masked or failure swallowed0.0% (0)0.0% (0)0.1% (1)Verification that cannot fail or misses target0.0% (0)0.0% (0)2.0% (21)Repository and test-suite hygiene violations0.0% (0)0.0% (0)0.0% (0)Deliverable or answer-format mismatch0.0% (0)0.0% (0)33.7% (363)Unsound analysis or validation methodology0.0% (0)0.0% (0)44.1% (475)

606 feedback observations (130 error-like) from 22 instances; 1,076 rubrics. Mismatch rubrics vs all feedback 97, vs error-like feedback 98.

The 20 induced buckets

idbucketdefinitionexamplessides
B01All tests passA real test-runner invocation (pytest, unittest, runtests.py) completes with every selected test passing, possibly with skips, xfails or warnings, and no failures or errors.OK — 42 tests passed, 3 skipped; '1 passed in 0.31s' with no failures reported; Ran 17 tests ... OK (skipped=2)feedback
B02Verified successful run or artifactA script, pipeline, or ad-hoc check runs to completion and produces positive evidence — PASS/SUCCESS markers, training metrics, or an output file whose format/contents are validated.'✓ All checks passed — issue is fixed' printed by the repro script; Training finished: CV AUC 0.912, submission.csv written with 10000 rows; Format check: columns, id coverage, value ranges all OKfeedback
B03Success with no pass/fail verdictThe command exits cleanly but yields only diagnostic dumps, a side-effect confirmation (file patched), silence, or truncated/empty output, so correctness must be inferred.Printed SQL, reprs and attribute dumps for inspection, no assertion made; 'Patch applied successfully' from a file-rewriting helper; rc=0 with empty output after grep/tail pipingfeedback
B04Environment not ready or not configuredThe run aborts before exercising any logic because a required package/tool is not installed or a framework was never bootstrapped (settings, env var, app registry, DB fixtures).ModuleNotFoundError: No module named 'pytest'; ImproperlyConfigured: settings are not configured before use; AppRegistryNotReady / missing DJANGO_SETTINGS_MODULEfeedback, rubric
B05Unresolved name, import path, or test selectorA specific target cannot be resolved: an imported symbol or attribute does not exist, a name is used without being defined/imported, the test label/path is unknown, or the CLI invocation itself is malformed.ImportError: cannot import name 'foo' from 'pkg.mod'; unittest.loader._FailedTest: module has no attribute 'TestX'; NameError: name 'helper' is not defined / unrecognized flag, usage errorfeedback, rubric
B06Source fails to parse, compile, or encodeExecution never happens (or dies while printing) because a file is syntactically invalid, indentation is broken by an edit, or output cannot be encoded by the console codec.SyntaxError: unterminated string literal in the heredoc script; IndentationError after the patch reparented a block; UnicodeEncodeError while printing '✓' to an ASCII stdoutfeedback, rubric
B07Timeout, hang, or resource exhaustionThe process produces no usable verdict because it exceeded the wall-clock limit, hung on an interactive prompt or infinite loop, or was killed for running out of memory/disk/shared memory.Command killed after 600s (rc=124) during model training; Process terminated with rc=137 'Killed' — out of memory; pip install exceeded the time limit; run dropped into pdb and never returnedfeedback
B08Missing, corrupt, or mismatched input dataThe run fails on the data or files it consumes: a path does not exist, a file cannot be parsed, a referenced column/key is absent, or array shapes and lengths do not line up.FileNotFoundError: submission.csv not found; ParserError: tokenizing data / cannot identify image file; KeyError: 'target' — column missing; shapes (1000,) and (998,) misalignedfeedback, rubric
B09API misuse or violated library preconditionA call is rejected by the callee's contract: unexpected/missing keyword, wrong arity or argument order, an unsupported option value, or data that violates a documented precondition of the estimator/helper.TypeError: __init__() got an unexpected keyword argument 'copy_x'; InvalidParameterError: 'auto' is no longer a supported value; ValueError: class labels must be contiguous / n_splits greater than members in a classfeedback, rubric
B10Uncaught runtime exception on a reachable pathExercised code raises an unhandled exception from an unguarded attribute, key, index, type, or None value, whether observed as a traceback or spotted as a guaranteed crash in the patch.AttributeError: 'NoneType' object has no attribute 'name' inside the library frame; IndexError from a hard-coded subscript into a variable-length list; TypeError: unsupported operand type(s) raised deep in the patched functionfeedback, rubric
B11Assertion or test failure (expected vs actual)Tests or scripts execute and report a concrete mismatch: a failing assertEqual/assert, a wrong exception or warning type, an image/golden comparison beyond tolerance, or a FAILED summary count.AssertionError: 'foo bar' != 'foo bar'; FAILED tests/test_api.py::test_roundtrip — 1 failed, 25 passed; DID NOT RAISE ValueError / image comparison RMS above tolerancefeedback
B12Silently wrong value or degenerate resultNothing raises, but the produced value is incorrect by construction or visibly wrong — inverted predicate, off-by-one/ordering slip, wrong operand or key, or a degenerate/chance-level output.Script exits 0 but prints 'expected 3, got 4 — FAIL'; Off-by-one boundary: interior cases right, edge index past the end; All predictions collapse to a single constant / accuracy at chance levelfeedback, rubric
B13Fix misses the defect (wrong site, no-op)The patch edits a sibling routine, unreached branch, comment/annotation, or a semantically equivalent expression — or nothing shippable at all — so the reported behavior is unchanged.Only comments, docstrings and formatting changed; the faulty constant is untouched; Guard added in a caller while the function the report blames is byte-identical; Change exists only in a scratch script or .patch file, never applied to the sourcerubric
B14Incomplete fix: sibling paths or dead optionOnly one of several parallel branches, call sites, keys, or duplicate definitions is repaired, or a newly accepted parameter/flag is never read, forwarded, or admitted by its gate.One dispatch arm fixed, the mirrored arm still reproduces the bug; Signature gains `force=True` but the body never reads it; Fix special-cases the reported value while the rest of the named class stays brokenrubric
B15Over-broad change or broken contractThe change reaches beyond the reported case — widened guards, altered defaults or output format, rewritten shared helpers, removed/renamed public names — regressing callers nobody complained about.Per-case option turned into a global default, changing already-correct inputs; Public keyword argument deleted, so existing callers raise TypeError; Adjacent working call sites rewritten, breaking behavior pinned by existing testsrubric
B16Symptom masked or failure swallowedThe complaint is silenced rather than fixed: broad except clauses, defensive clamps or fallbacks at the crash site, discarded exit codes, or deleted/weakened guards, validators, and tests.try/except around the failing call that only prints the error; hasattr/clamp guard added at the consumer while the mis-sized producer is untouched; Failing test deleted or its assertions stripped so the suite goes greenrubric
B17Verification that cannot fail or misses targetThe submitted check proves nothing: it prints instead of asserting, restates current behavior, uses stand-ins or the wrong copy of the code, never constructs the triggering condition, or dies in its own scaffolding.Script prints 'SUCCESS' unconditionally with no assert or non-zero exit; Test asserts against a locally re-implemented copy of the function; Repro builds only default inputs, so it passes identically on unfixed coderubric
B18Repository and test-suite hygiene violationsThe change set leaves scratch scripts, generated artifacts, or backup copies in the tree, drops files matching the test-discovery glob that mutate global state at import, or breaks lint/formatting conventions.test_check.py at repo root that calls sys.exit and os.chdir at import time; One-off source-patching script and .orig backup committed with the fix; Trailing newline removed in a lint-gated repo; unused imports left behindrubric
B19Deliverable or answer-format mismatchThe required artifact is missing, written to the wrong path, or its schema/row alignment/answer string deviates from the supplied template or literal answer format.submission.csv has renamed columns and a saved index, unlike sample_submission.csv; Predictions cover 900 of 1000 test ids, in a different order; Final answer hand-typed with extra quotes/brackets instead of the prescribed tokenrubric
B20Unsound analysis or validation methodologyThe result is not trustworthy because the prescribed spec/data was ignored or fabricated, rows or scope were silently changed, or quality is claimed without held-out validation, baselines, or sanity checks.Model fit on all rows and shipped with only an in-sample score; Referenced README/config never opened; bins and labels hard-coded from guesswork; dropna silently removes half the rows before the reported statistic is computedrubric
What this says. SWE-bench Verified: mismatch 73/100 against error-like feedback (85 against all feedback). Errors the agent meets far more often than the library addresses: Success with no pass/fail verdict (6% vs 0%), Verified successful run or artifact (6% vs 0%), Assertion or test failure (expected vs actual) (25% vs 3%), Environment not ready or not configured (18% vs 0%), Unresolved name, import path, or test selector (12% vs 2%). Rubric mass without a matching test-time signal: Silently wrong value or degenerate result (15% vs 1%), Incomplete fix: sibling paths or dead option (7% vs 0%), Fix misses the defect (wrong site, no-op) (21% vs 0%), Symptom masked or failure swallowed (6% vs 0%), Verification that cannot fail or misses target (15% vs 0%).
MLE-bench Lite: mismatch 98/100 against error-like feedback (97 against all feedback). Errors the agent meets far more often than the library addresses: Timeout, hang, or resource exhaustion (57% vs 0%), Uncaught runtime exception on a reachable path (18% vs 0%), Environment not ready or not configured (9% vs 0%). Rubric mass without a matching test-time signal: Silently wrong value or degenerate result (15% vs 2%), Deliverable or answer-format mismatch (34% vs 0%), Unsound analysis or validation methodology (44% vs 0%).

Test-time feedback vs rubric coverage (data-derived taxonomy)

How the buckets were made. Bottom-up from the corpus, not by hand. Corpus = every observation the interpreter agent received from a command the router classes as code execution (SWE-bench interp-c: 12,710 observations; MLE-bench mle-interp-c + mle-interp: 1,238; successes included) plus every rubric in the libraries the hybrid arms retrieve from (mined 1000 for SWE, on-policy all for MLE). Opus open-coded 48 stratified chunks of 50 items (feedback and rubric chunks, both benchmarks) into candidate categories, then merged all candidates into exactly 20 mutually exclusive buckets with definitions and examples. Haiku then labelled every observation and every rubric with one bucket (rubrics: plus an optional secondary). Code: analysis/taxonomy_induce.py.
Reading the charts. Three bars per bucket: all interpreter feedback (what the agent hears back, successes included), the error-like subset (non-zero return code or an error signature), and the rubrics (by the signal a violation would produce). Mismatch = total variation distance between two distributions in percentage points (0 = identical, 100 = disjoint).

SWE-bench Verified: share per bucket

all interpreter feedbackerror-like feedback onlyrubricsScript ran and reported success or results33.4% (4248)5.4% (188)0.0% (0)All targeted tests passed19.6% (2490)1.5% (51)0.0% (0)Silent success or diagnostic-only output16.8% (2131)2.4% (83)0.1% (1)Missing module, file, or unresolvable target7.1% (906)17.5% (609)0.6% (6)Assertion or check failed (expected vs actual)6.4% (818)19.9% (692)2.6% (26)Harness failure, timeout, or no usable output4.5% (569)15.9% (555)0.1% (1)Exception raised inside the code under test4.1% (524)14.2% (495)11.8% (118)Scratch script dies in its own scaffolding3.7% (466)11.7% (407)3.6% (36)Environment, config, or data prerequisite not ready2.8% (362)10.0% (349)0.5% (5)Runs clean but the result is wrong1.2% (152)0.5% (17)12.9% (129)Type, shape, or API contract mismatch0.2% (24)0.5% (19)3.3% (33)Edited source cannot parse, import, or resolve names0.1% (17)0.4% (15)3.8% (38)Incomplete fix: sibling paths left broken0.0% (2)0.1% (2)7.2% (72)Symptom suppressed instead of root cause fixed0.0% (1)0.0% (0)6.9% (69)No-op patch or fix at the wrong site0.0% (0)0.0% (0)22.3% (223)Over-broad change or collateral regression0.0% (0)0.0% (0)6.1% (61)Verification cannot fail or tests the wrong thing0.0% (0)0.0% (0)14.8% (148)Test and repository hygiene violations0.0% (0)0.0% (0)3.2% (32)Deliverable or format spec not followed0.0% (0)0.0% (0)0.1% (1)Unsound data handling or model validation0.0% (0)0.0% (0)0.1% (1)

12,710 feedback observations (3,482 error-like) from 500 instances; 1,000 rubrics. Mismatch rubrics vs all feedback 87, vs error-like feedback 79.

MLE-bench Lite: share per bucket

all interpreter feedbackerror-like feedback onlyrubricsScript ran and reported success or results57.8% (715)0.0% (0)0.0% (0)Harness failure, timeout, or no usable output21.2% (263)63.2% (251)0.0% (0)Silent success or diagnostic-only output7.9% (98)0.3% (1)0.0% (0)Exception raised inside the code under test4.7% (58)14.6% (58)0.1% (1)Missing module, file, or unresolvable target2.7% (34)8.6% (34)0.0% (0)Scratch script dies in its own scaffolding1.2% (15)2.8% (11)0.0% (0)Edited source cannot parse, import, or resolve names1.2% (15)3.8% (15)0.0% (0)Type, shape, or API contract mismatch1.1% (13)3.3% (13)0.0% (0)Environment, config, or data prerequisite not ready0.9% (11)2.8% (11)0.1% (1)Runs clean but the result is wrong0.9% (11)0.0% (0)4.7% (51)Assertion or check failed (expected vs actual)0.4% (5)0.8% (3)0.0% (0)All targeted tests passed0.0% (0)0.0% (0)0.0% (0)No-op patch or fix at the wrong site0.0% (0)0.0% (0)3.4% (37)Symptom suppressed instead of root cause fixed0.0% (0)0.0% (0)0.1% (1)Incomplete fix: sibling paths left broken0.0% (0)0.0% (0)0.1% (1)Over-broad change or collateral regression0.0% (0)0.0% (0)0.0% (0)Verification cannot fail or tests the wrong thing0.0% (0)0.0% (0)1.1% (12)Test and repository hygiene violations0.0% (0)0.0% (0)0.0% (0)Deliverable or format spec not followed0.0% (0)0.0% (0)37.1% (399)Unsound data handling or model validation0.0% (0)0.0% (0)53.3% (573)

1,238 feedback observations (397 error-like) from 44 instances; 1,076 rubrics. Mismatch rubrics vs all feedback 99, vs error-like feedback 100.

The 20 induced buckets

idbucketdefinitionexamplessides
B01Harness failure, timeout, or no usable outputThe command produced no interpretable result because the execution infrastructure errored, the process was killed by the time limit, or the captured output is empty/truncated with no verdict.rc=-1, 'Error executing command in GKE pod' — the command never really ran; Killed after the wall-clock limit (rc=124) mid-training, no final result; Non-zero exit with empty stdout/stderr, or only teardown noise in the captured tailfeedback
B02Missing module, file, or unresolvable targetExecution aborts before any project logic runs because an imported package/symbol, a native library, an invoked file, or the requested test selector cannot be resolved.ModuleNotFoundError: No module named 'pytest' / settings module not importable; ImportError: cannot import name X from Y; libGL.so.1 missing for a native extension; Runner reports unknown test id / 'can't open file' / zero tests collected / unrecognized flagfeedback
B03Environment, config, or data prerequisite not readyImports succeed but the run aborts because framework configuration, database/schema state, fixtures, or an expected input/intermediate file is missing or invalid.ImproperlyConfigured: settings are not configured / model not in INSTALLED_APPS; no such table / IntegrityError while setting up test data; FileNotFoundError for the dataset, checkpoint, or prior submission filefeedback
B04Exception raised inside the code under testThe reproduction or test runs far enough to trigger an unhandled traceback originating in the project's or library's own code, exposing the bug or an incomplete fix.Traceback ends inside django/db/models/... with a FieldError; Repro script raises TypeError from the library function being patched; Test ends in ERROR (exception escaped) rather than a failed assertionfeedback
B05Scratch script dies in its own scaffoldingThe agent's throwaway repro/verification script fails for reasons of its own making — guessed constructor or private API, undefined name, bad quoting, encoding of printed output — so the intended check never executes.NameError in the heredoc script; UnicodeEncodeError while printing ✓ characters; Script crashes building its fixture with a guessed API signature before reaching the call under test; Reproduction aborts in setup scaffolding, so the target code path is never reachedfeedback, rubric
B06Type, shape, or API contract mismatchA call boundary is violated: wrong keyword/arity for the installed API, missing column/key, wrong dtype or array shape, unseen category, or a value that may be None/short used without a guard.TypeError: __init__() got an unexpected keyword argument 'n_estimators'; KeyError on a column the loaded dataframe does not contain; shapes (n,) vs (n,m) mismatch; Patch assumes an attribute/return shape the library does not provide, so the first real use raisesfeedback, rubric
B07Assertion or check failed (expected vs actual)Code ran to completion but a test or hand-written check compared values and mismatched, including 'DID NOT RAISE', baseline-image diffs, and runs mixing passes with one or more failures.FAILED test_x - AssertionError: assert 3 == 4; DID NOT RAISE <class 'ValueError'>; Script prints 4 checks PASS and 1 FAIL / suite summary '2 failed, 30 passed'feedback
B08All targeted tests passedA real test-runner invocation completed with every selected test passing (possibly with skips, xfails, or warnings) and no failures or errors.OK (skipped=2) / '35 passed, 1 warning in 4.2s'; Ran 12 tests ... OK; Selected regression tests all green after the patchfeedback
B09Script ran and reported success or resultsAn ad-hoc verification script, pipeline, or validation run completes cleanly and reports the expected behavior, metrics, or a written output artifact.'All tests passed! ✓' from the hand-written scenario script; Training finished, CV AUC 0.87, submission.csv written with 5000 rows; Format check confirms columns, id set, value range and no NaNsfeedback
B10Silent success or diagnostic-only outputThe command exits cleanly but gives no pass/fail verdict: it only applied an edit, installed a package, compiled, or printed state for inspection.'File updated successfully' from a patch-applying script, nothing else; pip install finished with only root-user/upgrade notices; Prints dtypes, SQL, attribute dumps or source excerpts with no assertionfeedback
B11Runs clean but the result is wrongNo exception is raised, yet the produced value or behavior is incorrect — printed output contradicts expectations, or the patch's logic (inverted condition, off-by-one, wrong constant/operand, mis-ordered statements, loop/state bug) silently computes the wrong answer.Exit 0 but output shows 'expected 5, got 0' / NaNs still present in the predictions; Guard's comparison sense is flipped, so the interesting case returns the sentinel; Early return skips finalization; accumulator never populated, so an empty result is emittedfeedback, rubric
B12Edited source cannot parse, import, or resolve namesThe change leaves the file unusable at load time: syntax/indentation damage, leftover conflict markers, a missing import, a deleted public symbol, or a name read on a path where it is never bound.SyntaxError/IndentationError when compiling the file the agent just rewrote; Module references an alias or helper that is never imported -> NameError at import; UnboundLocalError: assignment lives in a sibling branch or after the userubric, feedback
B13No-op patch or fix at the wrong siteThe submission cannot change the reported behavior: cosmetic/equivalent rewrite, edits to a function or layer the reproducer never reaches, a new flag or parameter never read, or no source change at all.Only comments, formatting, or an equivalent expression changed; the faulty construct is byte-identical; Patch edits a neighbouring helper while the routine named in the traceback is untouched; New keyword argument accepted in the signature but never consulted in the bodyrubric
B14Symptom suppressed instead of root cause fixedThe failure is silenced rather than corrected — defensive guards or broadened except clauses at the consumer, errors swallowed and placeholders returned, or tests/harness edited or deleted to make the run green.hasattr/try-except wrapper hides the crash while the bad value is still produced upstream; Coercion failure absorbed, field dropped, caller sees a success-shaped wrong result; Pre-existing test module emptied or sleeps added so the suite passesrubric
B15Incomplete fix: sibling paths left brokenOnly the exact reproducer, one branch, one overload, or one of several co-reported symptoms is repaired while structurally identical call sites, subclasses, or near-miss inputs keep the old behavior.One dunder/override fixed, the analogous peers still reproduce the defect; Per-symptom edits that skip the shared helper both symptoms flow through; Second defect named in the issue never addressedrubric
B16Over-broad change or collateral regressionThe diff alters behavior beyond the reported scope — removing validation, options, or documented rules, widening shared defaults, changing messages/formats, or leaving out-of-diff callers, docs, and mirrored declarations inconsistent.Fix applied in the shared unconditional path, changing results for callers that were already correct; Public keyword parameter or precedence rule deleted with no replacement; Signature narrowed without updating other call sites, __all__, or the docstring that states the old valuerubric
B17Verification cannot fail or tests the wrong thingThe evidence offered for correctness is non-discriminating: prints instead of assertions, exceptions swallowed, expectations rewritten to current output, or the check exercises a different copy, entry point, or input than the reported case.Script prints 'ALL CHECKS PASSED' unconditionally; assert wrapped in a bare except; Re-implements the logic locally or imports the installed package instead of the edited source; Asserts guessed literals, or cites the already-green pre-existing suite as proofrubric
B18Test and repository hygiene violationsThe change set pollutes the tree or the suite: scratch files matching test discovery with import-time side effects, backups/generated artifacts committed, wall-clock sleeps, or unrestored global state.Root-level test_repro.py collected by pytest runs chdir/sys.exit at module scope; .orig/.bak duplicates, patcher scripts, or generated reports left in the repo; Test sleeps on real time or mutates a process-wide singleton without teardownrubric
B19Deliverable or format spec not followedThe required artifact is missing, written to an ad-hoc path, or its schema/row alignment/answer string deviates from the provided template or spec file, which was often never read.submission.csv has renamed/extra columns, an index column, or a row count that does not match the test set; Result reported only in prose; no file written at the expected path; README/sample-output that defines bins, labels, or the answer template was never openedrubric
B20Unsound data handling or model validationThe analysis or ML pipeline is methodologically invalid: fabricated or substituted inputs, wrong population/scope/unit of analysis, silent row loss, no held-out validation or baseline, off-metric selection, leakage, or a degenerate result accepted uncritically.Predictions shipped with only in-sample scores; no split, CV, or trivial baseline; Hard-coded/synthetic data used because the real file was not found; required filter or grouping ignored; Cluster count chosen by bare argmax yields singleton clusters, reported as successrubric
What this says. SWE-bench Verified: mismatch 79/100 against error-like feedback (87 against all feedback). Errors the agent meets far more often than the library addresses: Script ran and reported success or results (5% vs 0%), Missing module, file, or unresolvable target (17% vs 1%), Assertion or check failed (expected vs actual) (20% vs 3%), Harness failure, timeout, or no usable output (16% vs 0%), Scratch script dies in its own scaffolding (12% vs 4%), Environment, config, or data prerequisite not ready (10% vs 0%). Rubric mass without a matching test-time signal: Runs clean but the result is wrong (13% vs 0%), Incomplete fix: sibling paths left broken (7% vs 0%), Symptom suppressed instead of root cause fixed (7% vs 0%), No-op patch or fix at the wrong site (22% vs 0%), Over-broad change or collateral regression (6% vs 0%), Verification cannot fail or tests the wrong thing (15% vs 0%).
MLE-bench Lite: mismatch 100/100 against error-like feedback (99 against all feedback). Errors the agent meets far more often than the library addresses: Harness failure, timeout, or no usable output (63% vs 0%), Exception raised inside the code under test (15% vs 0%), Missing module, file, or unresolvable target (9% vs 0%). Rubric mass without a matching test-time signal: Deliverable or format spec not followed (37% vs 0%), Unsound data handling or model validation (53% vs 0%).