← back to the main dashboard · curation cost per rubric

General-lesson rubrics: does making the library applicable change anything?

Development on a SWE-bench slice. Sections 2–5 used 100 SWE-bench Verified instances as a dev slice (category drop v1→v2, channel variants) and then ran the picks on the other 400. The user holds the benchmarks out completely, so these are development artifacts, not results. Section 0 has the clean numbers.

0. Results (no SWE-bench in the development loop)

The 2026-09-09 canonical sweep, all 500 SWE-bench Verified instances, libraries mined and selected on the collection pool only. Everything below this section used a slice of SWE-bench for development and is kept for transparency, not as results.

armresolved / 500
interp-cinterpreter (ceiling)339
noexec-cfloor: no code execution, output withheld, no CWM297
haiku-judge-chybrid CWM, Haiku judge, mined 1000315
nolib-chybrid CWM, Opus judge, task criterion only309
scale100-chybrid CWM, Opus judge, mined 100303
random12-chybrid CWM, Opus judge, random 12302
haiku-nolib-chybrid CWM, Haiku judge, task criterion only298
scale10-chybrid CWM, Opus judge, mined 10297
scale1000-chybrid CWM, Opus judge, mined 1000285

Best without cheating: the interpreter. Best hybrid CWM without cheating: haiku-judge-c, below the interpreter. The general-v1 library with the canonical channel has one untuned run on 100 instances (61, equal to the task-criterion-only arm on the same 100, interpreter 68); its prompt was designed in a lab that scored prompt styles on 30 SWE-bench tasks, so even that run is not fully clean.

2026-09-10 overnight. Dev slice: 100 fixed SWE-bench Verified instances (data/subset100.json, seed 100); the other 400 instances are held for a clean test. Same Haiku 4.5 agent and never-execute hybrid CWM everywhere; arms differ only in the rubric library, the judge, and what the verdict text contains. Standard error at n=100 is about 5 points.

1. What changed in the library

2. Development runs on the 100-instance slice (not results)

general2-where-c70
interp-c68
haiku-judge-c64
general-listen-c63
general2-c62
nolib-c61
general-c61
noexec-c60
general2-listen-c59
general2-evid-c59
haiku-nolib-c58
scale100-c56
scale1000-c55

dark grey: interpreter (ceiling) · light grey: floor (execution commands never run, output withheld, no world model) · orange: hybrid CWM with the canonical titles-only verdict · yellow: channel variants (criterion check step, judge evidence rule, structured location).

armlibraryjudgeverdict channelresolved
general2-where-chybrid CWM, channel variantgeneral-v2 (181)Opustitles + structured location (path::symbol)70
interp-cinterpreter (ceiling)———68
haiku-judge-chybrid CWM, canonical channelmined 1000 (specific)Haiku 4.5titles64
general-listen-chybrid CWM, channel variantgeneral-v1 (191)Opustitles + criterion's check step63
general2-chybrid CWM, canonical channelgeneral-v2 (181)Opustitles62
nolib-chybrid CWM, canonical channeltask criterion onlyOpustitles61
general-chybrid CWM, canonical channelgeneral-v1 (191)Opustitles61
noexec-cfloor: no code execution, no CWM—— (no CWM)execution output withheld60
general2-listen-chybrid CWM, channel variantgeneral-v2 (181)Opustitles + criterion's check step59
general2-evid-chybrid CWM, channel variantgeneral-v2 (181)Opus, evidence ruletitles59
haiku-nolib-chybrid CWM, canonical channeltask criterion onlyHaiku 4.5titles58
scale100-chybrid CWM, canonical channelmined 100 (specific)Opustitles56
scale1000-chybrid CWM, canonical channelmined 1000 (specific)Opustitles55

3. Is the judge still handing out 1.00s, and does the agent listen?

From analysis/verdict_diagnosis.py. Calibration uses the last verdict before submit: how often a clean (1.00) or a red (<1.00) final verdict ended in a resolved instance. Listening: share of violated verdicts followed by an edit before the next verdict, and share of library violations still present at the next verdict.

armverdicts = 1.00final 1.00 → resolvedfinal <1.00 → resolvededit after violationedit after cleanviolations persistingechoes criterion
general2-where-c62%88% (n=66)33% (n=33)32%16%62%38%
general-listen-c27%96% (n=23)53% (n=77)27%20%68%20%
general2-c49%86% (n=58)27% (n=41)25%18%61%19%
general-c20%80% (n=25)56% (n=73)24%17%72%13%
general2-listen-c57%81% (n=57)31% (n=42)25%18%57%32%
general2-evid-c61%84% (n=61)21% (n=38)30%16%63%20%
Reading. The old libraries: 1.00 on 64–75% of verdicts, retrieved rubrics mostly not applicable. With general rubrics the judge has about 9 applicable criteria per verdict and is well calibrated at the end (general-v2: final 1.00 → 81–88% resolved, final red → 21–31%). The remaining gap to the interpreter is on the agent side: with titles alone or titles plus the criterion's check step, 57–72% of violations are still there at the next verdict. A verdict variant that carried a one-sentence reason from the judge did make the agent act (85/100 on the dev slice, 306/400 on the clean test) but was rejected by design: the channel is rubric→reward verdicts only, no free text from the judge. Its numbers are kept in docs/MORNING_REPORT_0908.md for the record and are not part of the results.

4. Which lessons carry signal (diagnosis only, never used for selection)

For the last verdict of each instance: how many times a category was flagged and how often those instances were not resolved, against the arm's unresolved base rate. From analysis/rubric_precision.py.

general-c · unresolved base rate 38%

categoryflagsunresolved when flagged
partial fix7583%
task criterion5879%
verification theatre4937%
reproducer ignored3534%
symptom patching1283%
wrong location1173%
scope creep667%
regression of existing contract580%
missed edge cases367%

general2-c · unresolved base rate 38%

categoryflagsunresolved when flagged
partial fix8680%
task criterion6891%
wrong location1974%
symptom patching1493%
regression of existing contract10100%
scope creep1080%
missed edge cases475%

Partial-fix, symptom-patching and wrong-location flags are right about 4 times in 5; the test-process categories flag at the base rate and were dropped in v2 on that principle (the metric does not score tests), not by fitting to these numbers.

5. The other 400 instances, for configurations chosen in section 2 (not clean)

armresolved / 400rate
interp-c (interpreter, ceiling)27167.8%
noexec-c (floor: no code execution, no CWM)23759.2%
general2-where-400 (general-v2, structured location)26566.2%

Efficiency on the same 400 (median wall clock per instance from pod timing records; agent cost from the agent's model stats; judge calls are graded verdicts, Opus fast mode, not priced here).

armmedian wall (min)agent $ / instanceagent API callsjudge calls
interp-c3.00.32680.0
noexec-c3.90.421010.0
nolib-c4.70.37909.4
scale1000-c5.90.368910.4
general2-where-4005.90.378710.0

Every hybrid arm costs about 2x wall clock (a judge verdict takes 8–15 s versus about 1 s for execution) and about 10 Opus calls per instance: no efficiency win.

5b. Oracle ceiling: is it retrieval or the rubrics? (diagnostic, 2026-09-11)

Per instance, Opus saw the issue, the gold patch and Haiku's failed patch, then either picked the best 12 of the 1000 library rubrics (perfect retrieval) or wrote 3 rubrics for that instance. Same agent, judge, channel and grading; only the rubric set changes. Gold patches touch nothing but the rubric choice, so these are ceilings, not method results. 303 instances = the 203 the floor fails + 100 it solves.

rubric sourceresolved / 303of 203 floor-failedof 100 floor-solved kept
floor (no execution, no CWM)1000100
normal retrieval, mined 10001324191
task criterion only1304189
ORACLE: best 12 of the 1000 (seed a)1293792
ORACLE: best 12 of the 1000 (seed b)1274186
ORACLE: best 4 of the 10001293792
interpreter (real execution)1607090
3 criteria written from the task text alone (no gold; method arm)1627191
ORACLE: 3 rubrics written with the gold patch in view (seed a)20510996
ORACLE: written (seed b)1939796
ORACLE: written 3 + library top-819410193

Single rubrics on the 70 instances only execution fixes (one rubric as the whole library, one real rollout per instance).

configurationresolved / 70 per run
no rubric, 5 repeated runs (chance baseline)26, 26, 28, 28, 32 (mean 28.0)
one library rubric (40 arms)min 23, median 29, max 36 (mean 29.1)
one rubric written for the instance (3 arms)32, 41, 47 (mean 40.0)
Reading. Perfect retrieval over the library changes nothing (127–129 vs 132 with ordinary retrieval; a single library rubric is worth about one instance over no rubric). Rubrics written for the instance, through the same channel, put the hybrid far above the interpreter and a single one is worth about twelve instances. The bottleneck is what the rubrics say, not how they are found; and the agent does act on a rubric that names the failure and where to look. Without the gold patch, three criteria the judge writes from the issue text alone reach interpreter parity on all 500 (338 vs 339).

6. Example rubrics from general-v2

Full text as the judge sees it. Browse every library on the rubric browser.

### Always update all required ancillary files (partial fix) - **Applies when**: Any bug-fix where the project’s contribution guide mandates companion changes such as release notes, documentation, test cases, or metadata. - **Failure it prevents**: A fix that corrects the code but is rejected by CI or maintainers because supporting files were left unchanged. - **Check (from the task and the diff alone, no execution)**: 1. Re-read the project’s “how to contribute” section for required non-code updates. 2. Inspect the current diff: verify at least one changelog / whats-new / docs file (or other mandated files) is touched. 3. Confirm the entry references the issue being fixed and follows the project’s formatting conventions. - **Not a violation when**: The project’s guidelines explicitly state that no ancillary update is needed for this class of change. applies to 22% of pool probes · mined from pandas-dev/pandas · curator o3
### Update all code paths and gates that implement the behavior, not just the first obvious condition (partial fix) - **Applies when**: A fix changes validation or enforcement logic that is controlled by multiple conditionals, entry points, or configuration flags. - **Failure it prevents**: Only adjusting one guard or site so the reproducer passes while other paths remain incorrect or previous guarantees are unintentionally removed. - **Check (from the task and the diff alone, no execution)**: 1. From the task, list the behaviors and flags that should influence the check. 2. In the diff, ensure every function/method that gates or performs that behavior is updated consistently (including decorated/overloaded/variant entry points). 3. Verify configuration options are threaded through and respected rather than removed. 4. Confirm no previously necessary guard was dropped without an equivalent check being added where responsibility properly belongs. - **Not a violation when**: The behavior is cleanly centralized and the change updates that single source of truth to cover all relevant cases. applies to 32% of pool probes · mined from python/mypy · curator gpt5
### Implement the full behavior change through the entire call path, not just a local tweak (partial fix) - **Applies when**: A fix or optimization depends on conditions that originate in one layer but must alter behavior across multiple functions or modules. - **Failure it prevents**: Making a small, localized change that does not actually alter the expensive or incorrect path in the reported scenario. - **Check (from the task and the diff alone, no execution)**: 1. From the task, note the precise condition under which behavior must change and the expensive action to be avoided. 2. In the diff, verify that any needed parameter/flag is introduced and passed through all involved functions to where the decision is made. 3. Confirm the diff actually bypasses or removes the costly call in that code path. 4. Ensure the observable behavior/return values remain consistent with the existing contract unless explicitly updated across all uses. - **Not a violation when**: The root cause is confined to one location and the single change demonstrably removes or corrects the faulty path without needing propagation. applies to 27% of pool probes · mined from iterative/dvc · curator gpt5
### Enumerate and handle every plausible edge-case input that can reach the modified code (missed edge cases) - **Applies when**: Fixing logic that branches on input values, types, or sentinels (e.g., nulls, special constants, option flags). - **Failure it prevents**: A patch that solves the reported example but still breaks on unconsidered value or flag combinations. - **Check (from the task and the diff alone, no execution)**: 1. From the issue description, list all variants of the problematic input (different null-likes, dtype flags, boolean options, etc.). 2. Trace each variant through the patched code; verify every branch now yields the intended behavior. 3. Look for early returns or conditionals that mention only some variants; confirm none are left untreated. - **Not a violation when**: The documented public API guarantees that only the specifically addressed variant can occur, making other edge cases impossible in practice. applies to 85% of pool probes · mined from pandas-dev/pandas · curator o3
### Fix the root cause instead of masking it with broad exception handling (symptom patching) - **Applies when**: A bug manifests through an exception or incorrect value and the proposed patch merely catches the exception, silences it, or overwrites the value without changing the logic that created the error. - **Failure it prevents**: Superficial “band-aid” fixes that leave the faulty path intact, letting other inputs still fail or causing silent regressions elsewhere. - **Check (from the task and the diff alone, no execution)**: 1. Look for newly-added try/except blocks, default-return branches, or silent passes that did not exist before. 2. Confirm no accompanying change removes or corrects the code that originally produced the bad state. 3. Ensure no new logical path asserts that the earlier failure “can be ignored” without adjusting upstream conditions. 4. Verify tests (if any) only cover the caught symptom and not the broader behavior. - **Not a violation when**: The exception handling is paired with a substantive upstream change that prevents the erroneous condition for all inputs and the catch block only provides a clear, documented fallback that maintains existing contracts. applies to 80% of pool probes · mined from pandas-dev/pandas · curator o3
### Fix the underlying logic rather than silencing or bypassing the error (symptom patching) - **Applies when**: A bug shows up as an exception, crash, or bad value, and the proposed patch simply catches the error or returns a default without changing the faulty computation. - **Failure it prevents**: Hiding the immediate symptom while leaving the real defect untouched, which resurfaces in other inputs or silently corrupts behavior. - **Check (from the task and the diff alone, no execution)**: 1. Spot new try/except blocks or guard clauses added exactly where the original traceback pointed. 2. See if the except body just `pass`, logs, or returns a placeholder without correcting state or inputs. 3. Confirm no deeper adjustment to the algorithm, data structures, or invariants accompanies the catch. - **Not a violation when**: The caught error is narrowly scoped and followed by a well-explained alternate implementation that preserves all functional guarantees. applies to 78% of pool probes · mined from iterative/dvc · curator o3