← back to the main dashboard · curation cost per rubric
General-lesson rubrics: does making the library applicable change anything?
Development on a SWE-bench slice. Sections 2–5 used 100 SWE-bench Verified instances as a dev slice (category drop v1→v2, channel variants) and then ran the picks on the other 400. The user holds the benchmarks out completely, so these are development artifacts, not results. Section 0 has the clean numbers.
0. Results (no SWE-bench in the development loop)
The 2026-09-09 canonical sweep, all 500 SWE-bench Verified instances, libraries mined and selected on the collection pool only. Everything below this section used a slice of SWE-bench for development and is kept for transparency, not as results.
| arm | | resolved / 500 |
interp-c | interpreter (ceiling) | 339 |
noexec-c | floor: no code execution, output withheld, no CWM | 297 |
haiku-judge-c | hybrid CWM, Haiku judge, mined 1000 | 315 |
nolib-c | hybrid CWM, Opus judge, task criterion only | 309 |
scale100-c | hybrid CWM, Opus judge, mined 100 | 303 |
random12-c | hybrid CWM, Opus judge, random 12 | 302 |
haiku-nolib-c | hybrid CWM, Haiku judge, task criterion only | 298 |
scale10-c | hybrid CWM, Opus judge, mined 10 | 297 |
scale1000-c | hybrid CWM, Opus judge, mined 1000 | 285 |
Best without cheating: the interpreter. Best hybrid CWM without cheating: haiku-judge-c, below the interpreter. The general-v1 library with the canonical channel has one untuned run on 100 instances (61, equal to the task-criterion-only arm on the same 100, interpreter 68); its prompt was designed in a lab that scored prompt styles on 30 SWE-bench tasks, so even that run is not fully clean.
2026-09-10 overnight. Dev slice: 100 fixed SWE-bench Verified instances (data/subset100.json, seed 100); the other 400 instances are held for a clean test. Same Haiku 4.5 agent and never-execute hybrid CWM everywhere; arms differ only in the rubric library, the judge, and what the verdict text contains. Standard error at n=100 is about 5 points.
1. What changed in the library
- Diagnosis on the old libraries: 91% of retrieved rubrics judged not applicable, reward 1.00 on 64–75% of verdicts, only 78 of 1000 rubrics ever fired: the rubrics were too specific to the failure they were mined from.
- New mining prompt (
mining/rubric_prompts.py, taxonomy prompt): each pool failure becomes a general lesson with applies when, failure it prevents, a check from task and diff alone, and an exemption. Curators o3 and GPT-5 (o3 gives the highest applicability, about 70%; Opus 18%, Sonnet 5 7–25%, Haiku 4.5 4–6%). - Validation without SWE-bench (
mining/mine_validated.py): 60 disjoint pool probes (30 wrong / 30 right patches from the SWE-gym / SWE-smith collection pool); keep rubrics applicable to at least 20% of probes; dedup at 0.90 → general-v1, 191 rubrics. No curator's rubrics separated wrong from right pool patches, so discrimination is not used for selection. - general-v2 (181): v1 minus the two test-process categories ("prove your fix with tests", "add a failing test first"), which produced 29% of all violations on the dev run while flagging at the unresolved base rate: SWE-bench scores the fix, not the agent's tests.
2. Development runs on the 100-instance slice (not results)
general2-where-c70
interp-c68
haiku-judge-c64
general-listen-c63
general2-c62
nolib-c61
general-c61
noexec-c60
general2-listen-c59
general2-evid-c59
haiku-nolib-c58
scale100-c56
scale1000-c55
dark grey: interpreter (ceiling) · light grey: floor (execution commands never run, output withheld, no world model) · orange: hybrid CWM with the canonical titles-only verdict · yellow: channel variants (criterion check step, judge evidence rule, structured location).
| arm | library | judge | verdict channel | resolved |
general2-where-chybrid CWM, channel variant | general-v2 (181) | Opus | titles + structured location (path::symbol) | 70 |
interp-cinterpreter (ceiling) | — | — | — | 68 |
haiku-judge-chybrid CWM, canonical channel | mined 1000 (specific) | Haiku 4.5 | titles | 64 |
general-listen-chybrid CWM, channel variant | general-v1 (191) | Opus | titles + criterion's check step | 63 |
general2-chybrid CWM, canonical channel | general-v2 (181) | Opus | titles | 62 |
nolib-chybrid CWM, canonical channel | task criterion only | Opus | titles | 61 |
general-chybrid CWM, canonical channel | general-v1 (191) | Opus | titles | 61 |
noexec-cfloor: no code execution, no CWM | — | — (no CWM) | execution output withheld | 60 |
general2-listen-chybrid CWM, channel variant | general-v2 (181) | Opus | titles + criterion's check step | 59 |
general2-evid-chybrid CWM, channel variant | general-v2 (181) | Opus, evidence rule | titles | 59 |
haiku-nolib-chybrid CWM, canonical channel | task criterion only | Haiku 4.5 | titles | 58 |
scale100-chybrid CWM, canonical channel | mined 100 (specific) | Opus | titles | 56 |
scale1000-chybrid CWM, canonical channel | mined 1000 (specific) | Opus | titles | 55 |
3. Is the judge still handing out 1.00s, and does the agent listen?
From analysis/verdict_diagnosis.py. Calibration uses the last verdict before submit: how often a clean (1.00) or a red (<1.00) final verdict ended in a resolved instance. Listening: share of violated verdicts followed by an edit before the next verdict, and share of library violations still present at the next verdict.
| arm | verdicts = 1.00 | final 1.00 → resolved | final <1.00 → resolved | edit after violation | edit after clean | violations persisting | echoes criterion |
general2-where-c | 62% | 88% (n=66) | 33% (n=33) | 32% | 16% | 62% | 38% |
general-listen-c | 27% | 96% (n=23) | 53% (n=77) | 27% | 20% | 68% | 20% |
general2-c | 49% | 86% (n=58) | 27% (n=41) | 25% | 18% | 61% | 19% |
general-c | 20% | 80% (n=25) | 56% (n=73) | 24% | 17% | 72% | 13% |
general2-listen-c | 57% | 81% (n=57) | 31% (n=42) | 25% | 18% | 57% | 32% |
general2-evid-c | 61% | 84% (n=61) | 21% (n=38) | 30% | 16% | 63% | 20% |
Reading. The old libraries: 1.00 on 64–75% of verdicts, retrieved rubrics mostly not applicable. With general rubrics the judge has about 9 applicable criteria per verdict and is well calibrated at the end (general-v2: final 1.00 → 81–88% resolved, final red → 21–31%). The remaining gap to the interpreter is on the agent side: with titles alone or titles plus the criterion's check step, 57–72% of violations are still there at the next verdict. A verdict variant that carried a one-sentence reason from the judge did make the agent act (85/100 on the dev slice, 306/400 on the clean test) but was rejected by design: the channel is rubric→reward verdicts only, no free text from the judge. Its numbers are kept in docs/MORNING_REPORT_0908.md for the record and are not part of the results.
4. Which lessons carry signal (diagnosis only, never used for selection)
For the last verdict of each instance: how many times a category was flagged and how often those instances were not resolved, against the arm's unresolved base rate. From analysis/rubric_precision.py.
general-c · unresolved base rate 38%
| category | flags | unresolved when flagged |
| partial fix | 75 | 83% |
| task criterion | 58 | 79% |
| verification theatre | 49 | 37% |
| reproducer ignored | 35 | 34% |
| symptom patching | 12 | 83% |
| wrong location | 11 | 73% |
| scope creep | 6 | 67% |
| regression of existing contract | 5 | 80% |
| missed edge cases | 3 | 67% |
general2-c · unresolved base rate 38%
| category | flags | unresolved when flagged |
| partial fix | 86 | 80% |
| task criterion | 68 | 91% |
| wrong location | 19 | 74% |
| symptom patching | 14 | 93% |
| regression of existing contract | 10 | 100% |
| scope creep | 10 | 80% |
| missed edge cases | 4 | 75% |
Partial-fix, symptom-patching and wrong-location flags are right about 4 times in 5; the test-process categories flag at the base rate and were dropped in v2 on that principle (the metric does not score tests), not by fitting to these numbers.
5. The other 400 instances, for configurations chosen in section 2 (not clean)
| arm | resolved / 400 | rate |
interp-c (interpreter, ceiling) | 271 | 67.8% |
noexec-c (floor: no code execution, no CWM) | 237 | 59.2% |
general2-where-400 (general-v2, structured location) | 265 | 66.2% |
Efficiency on the same 400 (median wall clock per instance from pod timing records; agent cost from the agent's model stats; judge calls are graded verdicts, Opus fast mode, not priced here).
| arm | median wall (min) | agent $ / instance | agent API calls | judge calls |
interp-c | 3.0 | 0.32 | 68 | 0.0 |
noexec-c | 3.9 | 0.42 | 101 | 0.0 |
nolib-c | 4.7 | 0.37 | 90 | 9.4 |
scale1000-c | 5.9 | 0.36 | 89 | 10.4 |
general2-where-400 | 5.9 | 0.37 | 87 | 10.0 |
Every hybrid arm costs about 2x wall clock (a judge verdict takes 8–15 s versus about 1 s for execution) and about 10 Opus calls per instance: no efficiency win.
5b. Oracle ceiling: is it retrieval or the rubrics? (diagnostic, 2026-09-11)
Per instance, Opus saw the issue, the gold patch and Haiku's failed patch, then either picked the best 12 of the 1000 library rubrics (perfect retrieval) or wrote 3 rubrics for that instance. Same agent, judge, channel and grading; only the rubric set changes. Gold patches touch nothing but the rubric choice, so these are ceilings, not method results. 303 instances = the 203 the floor fails + 100 it solves.
| rubric source | resolved / 303 | of 203 floor-failed | of 100 floor-solved kept |
| floor (no execution, no CWM) | 100 | 0 | 100 |
| normal retrieval, mined 1000 | 132 | 41 | 91 |
| task criterion only | 130 | 41 | 89 |
| ORACLE: best 12 of the 1000 (seed a) | 129 | 37 | 92 |
| ORACLE: best 12 of the 1000 (seed b) | 127 | 41 | 86 |
| ORACLE: best 4 of the 1000 | 129 | 37 | 92 |
| interpreter (real execution) | 160 | 70 | 90 |
| 3 criteria written from the task text alone (no gold; method arm) | 162 | 71 | 91 |
| ORACLE: 3 rubrics written with the gold patch in view (seed a) | 205 | 109 | 96 |
| ORACLE: written (seed b) | 193 | 97 | 96 |
| ORACLE: written 3 + library top-8 | 194 | 101 | 93 |
Single rubrics on the 70 instances only execution fixes (one rubric as the whole library, one real rollout per instance).
| configuration | resolved / 70 per run |
|---|
| no rubric, 5 repeated runs (chance baseline) | 26, 26, 28, 28, 32 (mean 28.0) |
| one library rubric (40 arms) | min 23, median 29, max 36 (mean 29.1) |
| one rubric written for the instance (3 arms) | 32, 41, 47 (mean 40.0) |
Reading. Perfect retrieval over the library changes nothing (127–129 vs 132 with ordinary retrieval; a single library rubric is worth about one instance over no rubric). Rubrics written for the instance, through the same channel, put the hybrid far above the interpreter and a single one is worth about twelve instances. The bottleneck is what the rubrics say, not how they are found; and the agent does act on a rubric that names the failure and where to look. Without the gold patch, three criteria the judge writes from the issue text alone reach interpreter parity on all 500 (338 vs 339).
6. Example rubrics from general-v2
Full text as the judge sees it. Browse every library on the rubric browser.
### Always update all required ancillary files (partial fix)
- **Applies when**: Any bug-fix where the project’s contribution guide mandates companion changes such as release notes, documentation, test cases, or metadata.
- **Failure it prevents**: A fix that corrects the code but is rejected by CI or maintainers because supporting files were left unchanged.
- **Check (from the task and the diff alone, no execution)**:
1. Re-read the project’s “how to contribute” section for required non-code updates.
2. Inspect the current diff: verify at least one changelog / whats-new / docs file (or other mandated files) is touched.
3. Confirm the entry references the issue being fixed and follows the project’s formatting conventions.
- **Not a violation when**: The project’s guidelines explicitly state that no ancillary update is needed for this class of change.
applies to 22% of pool probes · mined from pandas-dev/pandas · curator o3
### Update all code paths and gates that implement the behavior, not just the first obvious condition (partial fix)
- **Applies when**: A fix changes validation or enforcement logic that is controlled by multiple conditionals, entry points, or configuration flags.
- **Failure it prevents**: Only adjusting one guard or site so the reproducer passes while other paths remain incorrect or previous guarantees are unintentionally removed.
- **Check (from the task and the diff alone, no execution)**: 1. From the task, list the behaviors and flags that should influence the check. 2. In the diff, ensure every function/method that gates or performs that behavior is updated consistently (including decorated/overloaded/variant entry points). 3. Verify configuration options are threaded through and respected rather than removed. 4. Confirm no previously necessary guard was dropped without an equivalent check being added where responsibility properly belongs.
- **Not a violation when**: The behavior is cleanly centralized and the change updates that single source of truth to cover all relevant cases.
applies to 32% of pool probes · mined from python/mypy · curator gpt5
### Implement the full behavior change through the entire call path, not just a local tweak (partial fix)
- **Applies when**: A fix or optimization depends on conditions that originate in one layer but must alter behavior across multiple functions or modules.
- **Failure it prevents**: Making a small, localized change that does not actually alter the expensive or incorrect path in the reported scenario.
- **Check (from the task and the diff alone, no execution)**: 1. From the task, note the precise condition under which behavior must change and the expensive action to be avoided. 2. In the diff, verify that any needed parameter/flag is introduced and passed through all involved functions to where the decision is made. 3. Confirm the diff actually bypasses or removes the costly call in that code path. 4. Ensure the observable behavior/return values remain consistent with the existing contract unless explicitly updated across all uses.
- **Not a violation when**: The root cause is confined to one location and the single change demonstrably removes or corrects the faulty path without needing propagation.
applies to 27% of pool probes · mined from iterative/dvc · curator gpt5
### Enumerate and handle every plausible edge-case input that can reach the modified code (missed edge cases)
- **Applies when**: Fixing logic that branches on input values, types, or sentinels (e.g., nulls, special constants, option flags).
- **Failure it prevents**: A patch that solves the reported example but still breaks on unconsidered value or flag combinations.
- **Check (from the task and the diff alone, no execution)**:
1. From the issue description, list all variants of the problematic input (different null-likes, dtype flags, boolean options, etc.).
2. Trace each variant through the patched code; verify every branch now yields the intended behavior.
3. Look for early returns or conditionals that mention only some variants; confirm none are left untreated.
- **Not a violation when**: The documented public API guarantees that only the specifically addressed variant can occur, making other edge cases impossible in practice.
applies to 85% of pool probes · mined from pandas-dev/pandas · curator o3
### Fix the root cause instead of masking it with broad exception handling (symptom patching)
- **Applies when**: A bug manifests through an exception or incorrect value and the proposed patch merely catches the exception, silences it, or overwrites the value without changing the logic that created the error.
- **Failure it prevents**: Superficial “band-aid” fixes that leave the faulty path intact, letting other inputs still fail or causing silent regressions elsewhere.
- **Check (from the task and the diff alone, no execution)**:
1. Look for newly-added try/except blocks, default-return branches, or silent passes that did not exist before.
2. Confirm no accompanying change removes or corrects the code that originally produced the bad state.
3. Ensure no new logical path asserts that the earlier failure “can be ignored” without adjusting upstream conditions.
4. Verify tests (if any) only cover the caught symptom and not the broader behavior.
- **Not a violation when**: The exception handling is paired with a substantive upstream change that prevents the erroneous condition for all inputs and the catch block only provides a clear, documented fallback that maintains existing contracts.
applies to 80% of pool probes · mined from pandas-dev/pandas · curator o3
### Fix the underlying logic rather than silencing or bypassing the error (symptom patching)
- **Applies when**: A bug shows up as an exception, crash, or bad value, and the proposed patch simply catches the error or returns a default without changing the faulty computation.
- **Failure it prevents**: Hiding the immediate symptom while leaving the real defect untouched, which resurfaces in other inputs or silently corrupts behavior.
- **Check (from the task and the diff alone, no execution)**:
1. Spot new try/except blocks or guard clauses added exactly where the original traceback pointed.
2. See if the except body just `pass`, logs, or returns a placeholder without correcting state or inputs.
3. Confirm no deeper adjustment to the algorithm, data structures, or invariants accompanies the catch.
- **Not a violation when**: The caught error is narrowly scoped and followed by a well-explained alternate implementation that preserves all functional guarantees.
applies to 78% of pool probes · mined from iterative/dvc · curator o3