Why the hybrid CWM does not beat the interpreter: the diagnosis
SWE-bench Verified, all 500 instances held out · each section: the experiment, the metric, the result, one paragraph on what it means · numbers recomputed from the run artifacts on 2026-09-16 09:21
▸How each method works
Interpreter — every command runs; the agent sees real output. The ceiling.Floor, no execution — reads and edits are real; runs return nothing. The floor.Hybrid CWM — reads go to the interpreter; runs go to the world model, which grades the diff against retrieved rubrics and returns reward + rubric text.Route-by-cost — only expensive execution goes to the world model; everything else runs for real.
The Sonnet 5 arms are these same four methods with Sonnet 5 in the agent box; judge, library, channel and routing are unchanged.
Setup. Every arm runs the same agent harness on the same 500 instances, graded by the official FAIL_TO_PASS / PASS_TO_PASS tests. Interpreter: every command runs in the sandbox. Floor: commands the router classes as code execution never run and return nothing. Hybrid CWM: those commands never run; the judge grades the agent's current git diff against the rubrics retrieved for the task (mined 1000 = the 1,000-rubric library mined from failed agent patches on SWE-Gym / SWE-smith) and the agent receives Predicted reward: X (N applicable criteria) followed by each applicable rubric's full text with its own [reward 0|1]. Preamble: a fixed harness sentence in front of that feedback saying it is reward-model feedback on the current change and that every [reward 0] criterion must be addressed. Haiku judge: same arm with Haiku 4.5 grading instead of Opus. Route-by-cost: only expensive execution (suite-level tests, training, benchmarks) goes to the CWM; everything else runs for real.
Haiku 4.5 as the coding agent: resolved instances, all 500
A second identical interpreter run gave 355 and differs from this one on 70 instances, so differences under about 10 are run-to-run noise. Route-by-cost consulted the CWM on 72 of 500 instances (115 verdicts in total): it is an interpreter replicate.
Where each command went, Haiku 4.5 agent (inner ring: destination; outer ring: command type by the router's classifier)
Destination is read from what came back: a world-model reply = CWM; "<execution output withheld>" = the floor's empty answer; anything else ran in the sandbox. In the hybrid arms every code-execution command goes to the CWM; in route-by-cost only the expensive ones do.
interpreter · 33,366 commands, 67 per instancefloor: no execution, no CWM · 50,377 commands, 101 per instancehybrid CWM · mined 1000 · 39,685 commands, 79 per instancehybrid CWM · mined 1000 + preamble · 40,569 commands, 81 per instancehybrid CWM · Haiku judge · mined 1000 · 40,849 commands, 82 per instanceroute-by-cost · 31,751 commands, 64 per instance
Sonnet 5 as the coding agent: the same six arms, only the agent model changed
Same judge, library, channel, preamble and routing rules as the Haiku arms above.
Where each command went, Sonnet 5 agent
interpreter · 12,465 commands, 25 per instancefloor: no execution, no CWM · 33,902 commands, 68 per instancehybrid CWM · mined 1000 · 28,280 commands, 57 per instancehybrid CWM · mined 1000 + preamble · 24,814 commands, 50 per instancehybrid CWM · Haiku judge · mined 1000 · 28,047 commands, 56 per instanceroute-by-cost · 12,943 commands, 26 per instance
1Was the agent ignoring the feedback? Measuring listening, then the preamble fix
Metric. A red verdict is a world-model reply in which at least one applicable rubric got [reward 0] (predicted reward below 1.00); a clean verdict has every rubric at [reward 1]. Edit after red = of all red verdicts in an arm, the share where the agent issued at least one source-editing command (sed -i, heredoc or redirect that writes a file, patch, git apply, mv/cp) before its next routed command. If the agent goes straight to re-running tests or reading files, the red was ignored. Denominators are printed under each chart.
The fix. The preamble: fixed harness text (no judge language) in front of every verdict stating that the output is reward-model feedback on the current diff, that [reward 0] criteria are violated and must be addressed before running tests again, and that the command was not executed. Shown verbatim below; nothing else changed between the paired arms.
The preamble, verbatim
before the verdict:
[Coding world model] Your command was NOT executed. A reward model graded your current change (git diff) against review criteria. Below: the predicted reward, then every applicable criterion with its reward (1 = satisfied, 0 = violated).
after the last criterion:
Treat each [reward 0] criterion as feedback on your code: locate where it applies in this repository (grep, cat and sed -n run for real) and edit the code to satisfy it, unless it clearly does not apply to this task. Running code again only re-grades the same change; edit first.
Fixed harness text, identical for every task and every verdict; it contains no judge output. The full tool result the agent sees is below.
What the agent receives (verbatim, one routed command in astropy__astropy-13033, mined 1000 + preamble)
[Coding world model] Your command was NOT executed. A reward model graded your current change (git diff) against review criteria. Below: the predicted reward, then every applicable criterion with its reward (1 = satisfied, 0 = violated).
Predicted reward: 0.00 (3 applicable criteria)
[reward 0]
### The change does not implement the behavior the task requires
- **Applies when**: always -- this criterion is derived from the task statement itself.
- **Pattern**: The diff does not make the behavior described in the task actually happen:
it edits the wrong location, changes something adjacent to the reported symptom, only
adds tests/scripts/logging, or handles a different case than the one reported. Apply the
diff mentally and ask whether the exact symptom described in the task would still occur.
- **Detection procedure**:
1. From the task statement, state the concrete before/after behavior being requested
(the call that misbehaves and what it should do instead). [reads: task]
2. Locate the lines in the diff that would change that behavior. If no hunk lies on the
execution path of the reported call, the criterion fires. [reads: code]
3. If a hunk is on that path, check that it produces the requested behavior for the
reported input -- not merely that it removes an exception or renames something.
- **Discriminator**: fires when the reported symptom survives the patch; does not fire when
the diff demonstrably changes the reported behavior to what the task asks for.
- **Consequence**: the task's own failing test still fails; the submission is scored wrong.
[reward 0]
### Library behavior changed but verified only by throwaway scripts outside the project's test directory
- **Applies when**: `task|code`: the task asks for a source change in a repository whose static facts list a dedicated test directory containing a test module for the source file being modified.
- **Pattern**: The program edits the library source and adds verification as new top-level scripts (print/assert `__main__` runners, patch-applier scripts that rewrite source files by string replacement, summary markdown), while adding no case to the repository's existing test module for the changed code.
- **Detection procedure**:
1. List the files the program creates or modifies and identify which are under the repository's test directory. [reads: code]
2. Compare against the static facts repo tree: does a test module exist that corresponds to the modified source module? [reads: static facts — repo tree]
3. It fires when the modified source file has a corresponding test module in that directory, none of those test files are touched, and the new files are repo-root scripts containing module-level statements or `if __name__ == "__main__":` runners, or scripts that open the source file and `content.replace(...)`/write it back. [reads: code]
- **Counter-example**: a program that adds or extends test functions inside the repository's existing test module (even if it also leaves one scratch script), or a repository whose static facts show no test directory at all.
- **Discriminator**: the new behavior has zero coverage in the collected suite — the only assertions live in files that the project's test discovery configuration does not target, so a green suite proves nothing about the change.
- **Consequence**: regressions introduced by the edit (removed checks, changed types) go undetected and hidden/maintainer tests for the changed behavior fail; additionally, root-level files named `test_*.py` that execute code at import are collected when pytest is run from the repository root, producing collection-time `AssertionError`/`ImportError`/`SyntaxError` and non-zero exit. Explains the discrepancy between a fully passing run and a semantically wrong patch; the remaining risk comes from the substantive defects in the edit itself.
- **Evidence**: the source module was edited while its sibling test module in the project's test directory was left untouched; verification consisted of `fix.py`/`fix2.py`/`fix3.py`/`fix4.py` string-replacement patchers and several root-level `test_*.py` print-and-assert scripts, and the reported `2386 passed` covered none of the new behavior.
[reward 0]
### Bug fix gated on a data-dependent label lookup instead of the declared argument
- **Applies when**: `task`: the task is a bug report with a reproducer in which a library API raises on input the reporter considers valid, and `code`: the patch adds a new conditional in front of the previously failing code path.
- **Pattern**: The repair is implemented as a runtime guess — the new branch inspects the *contents* of the object being operated on (e.g. builds sets of two competing label namespaces and intersects them with the user-supplied keys) to decide which of two incompatible semantics to apply — instead of deciding from the explicit argument that names the semantics. The corrected path therefore fires only for inputs where the two namespaces happen not to overlap; every other input silently falls through to the original faulty code.
- **Detection procedure**:
1. Locate the conditional block added immediately before the statement that produced the reported failure (the transpose, the re-dispatch, the lookup that raised) [reads: code]
2. Compare the condition's inputs against the argument the task's reproducer actually passes: does the condition read only the explicit flag/argument, or does it also read runtime containers of the operand (its index labels, column labels, key sets, dtypes) [reads: task statement (the reproducer call and its arguments) + code]
3. Check the fall-through: is the `else`/unguarded continuation byte-for-byte the original failing path, and can an input of the same shape as the reproducer fail the guard — e.g. one key present in *both* namespaces, or only a subset of keys matching, making the intersection test false [reads: code]
- **Counter-example**: A branch that dispatches on the explicit parameter (`axis`, `how`, a type check on `func`) and keeps the old path only for genuinely different argument kinds, so every input matching the reported scenario takes the new path regardless of what labels the data carries.
- **Discriminator**: In the failing case, whether the fix applies depends on incidental overlap between two label collections in the user's data; in the safe case the dispatch is decided entirely by arguments and types, independent of the data values.
- **Consequence**: The exact reproducer passes while near-identical hidden cases still raise the original exception (`KeyError`, or the library's lookup error) — e.g. a frame whose index labels coincide with its column labels, or a dict covering only some labels. Predict partial credit at best: reproducer-only tests pass, parametrized/edge-case tests for the same API fail.
- **Evidence**: `func_keys = set(func.keys()); ... if len(func_keys) == len(keys_in_columns) and len(keys_in_index) == 0: return self.transform_dict_like(func)` — a guess between "keys are columns" and "keys are row labels" that reverts to the old raising path whenever any key also appears in the other axis.
Treat each [reward 0] criterion as feedback on your code: locate where it applies in this repository (grep, cat and sed -n run for real) and edit the code to satisfy it, unless it clearly does not apply to this task. Running code again only re-grades the same change; edit first.
The command the agent issued was never executed; this text is the entire tool result. In the case browser every verdict expands to this full text.
Listening before and after the preamble
Red verdicts: scale1000-r: 1255; benchpre-mined1000: 912. 'Still red at submit' = of the instances that ever saw a red verdict, the share whose last verdict before submit was still red. The agent's next message names the violated criterion (two or more of its content words) after 2% vs 18% of red verdicts, without vs with the preamble.
Resolved, all 500, before and after the preamble
Interpreter 339; floor 297.
What this says. Before the preamble the agent edited after a red verdict 35% of the time (mined 1000); with it, 53%, and fewer reds survive to submit. Resolve moves by +2: the agent now acts on the flags and the library still stays 24 below the interpreter.
2Is Haiku the bottleneck? Sonnet 5 as the coding agent
Setup. Swap the coding agent for Sonnet 5 and change nothing else: same judge, same mined-1000 library, same channel, preamble and routing rules. Run the same six arms on all 500. If the hybrid closes the gap to Sonnet's own interpreter, Haiku's ability to act on a flag was the limit; if the gap persists or widens, the feedback is.
Metric.Judge on the final patch: the last verdict before submit is compared with the official grade of that submitted patch. Sensitivity = share of wrong final patches the judge flagged red; false-positive rate = share of right (resolved) final patches it flagged red; precision = share of red final verdicts that were on a wrong patch.
Resolved, all 500, by coding agent
The same Opus judge on each agent's final patch (mined 1000 + preamble)
Haiku: 182 wrong / 309 right final patches; Sonnet: 137 / 352.
What this says. Haiku's hybrid (mined 1000 + preamble) is 24 below its interpreter; Sonnet's is 71 below its interpreter and 34 relative to its own floor. The judge flags 22% of Sonnet's wrong patches vs 55% of Haiku's; 43% of the flags Sonnet receives are on correct patches. Sonnet edits after a red 36% of the time (Haiku 53%).
3Do the rubrics apply to unseen repositories, and are they right when they fire?
Setup. A judge screen on a repository-disjoint validation pool: 97 agent patches from SWE-Gym / SWE-smith tasks in repositories never used for mining, each with its real test outcome (60 wrong, 37 right). For every rubric × probe the Opus judge answers three questions from the task text and the diff alone: does this rubric apply here, is it violated, and where. The contracts differ only in the prompt used to write rubrics from the same failed agent patches.
Metrics.Applicability = share of the 97 probes the judge says the rubric applies to. Fires on wrong patches = of the wrong probes it applies to, the share it marks violated (a true flag). Fires on right patches = of the right probes it applies to, the share it marks violated (a false flag). Localized = share of firings that name a file/function. A rubric counts as transferable if it applies to at least 15% of probes, fires on at least 12% of wrong and at most 5% of right patches, and localizes at least half the time. Specific is the original mined-1000 style ('describe the failure pattern you see'); relational writes a generic check anchored on an artifact of the task statement (traceback, expected output, listed variants, named symbol, reproducer) plus one relation to the diff.
Applicability on unseen repositories (mean share of the 97 probes a rubric applies to), with the count passing all four transfer thresholds
On the benchmark. For the library arms, the final verdict's library flags are compared with the official grade of the submitted patch: precision = share of library red flags that were on a wrong patch, against the base rate of wrong final patches in that arm.
Precision of library flags on the final patch (official grade as truth)
What this says. The specific style applies to almost no task outside its source repository, which is why the mined-1000 library rarely fires. The relational contract applies to about half of unseen tasks and fires with a location, so applicability is fixed. On the benchmark a library flag is right about half the time at a base rate around a third: it carries a type of mistake but not where the patch is wrong, and every library arm stays below the interpreter.
4Is it retrieval? Oracle selection with the gold patch
Setup. The gold patch is the reference human fix of the SWE-bench instance (the merged pull request); it is used here only as a diagnostic, never in a method arm. On 303 instances (all floor failures plus a sample of floor successes) Opus saw the issue, the gold patch and Haiku's failed patch, then either picked the best 12 (or 4) of the 1000 library rubrics (perfect retrieval) or wrote 3 rubrics for that instance. The chosen rubrics were installed as the whole library for one rollout with the same agent, judge and channel.
Resolved on the 303 oracle instances
What this says. Perfect retrieval from the library changes nothing (within a few instances of ordinary retrieval). Rubrics written for the instance through the same channel beat the interpreter. The channel can carry what the agent needs; the library's content does not contain it.
5Is it the judge? Opus vs Haiku as the world model
Setup and metric. Same mined-1000 library and Haiku agent, judge swapped. Calibration of the last verdict before submit: of the instances whose last verdict was clean, how many resolved; same for red. A calibrated judge has a large gap between the two.
Calibration of the last verdict before submit
Opus: 343 clean / 150 red final verdicts, resolved 313; Haiku: 359 / 131, resolved 307.
What this says. Opus's verdicts are informative (clean → about 80% resolved, red → about 20%) and Haiku's barely are, yet both arms land near the floor plus a little. Judge exemplars (4 executed pool patches with real outcomes in the judge prompt) raise Opus's precision further with no reliable change in resolve. Judge quality is not the binding constraint while a flag carries no location.
6Case browser: trajectories behind the numbers
How to read a trace. Click a baseline to open its examples, then an instance to open the trace. Each row is one agent command (edit, run, read), the sandbox's real output (grey), or the world model's answer. Verdict rows show the predicted reward and each applicable rubric's [reward]; expand 'exact text the agent received' for the full preamble and rubric bodies. Chips give the official grade of this arm's final patch and of the interpreter and floor on the same instance.
Sources
Numbers recomputed at build time from reports/ and results/*.traj.json by make_diagnosis_page.py. Listening uses the same definition as analysis/verdict_diagnosis.py; calibration uses the last verdict before submit against the official grade. Narrative and per-arm logs: docs/DIAGNOSIS_0912.md, docs/DIAGNOSIS_0912_cases.md, docs/MORNING_REPORT_0908.md in the agent-cwm repository. Companion pages: SWE-bench reward-channel dashboard, general-lesson rubrics and oracle tables, rubric browser.