CAUSAL PRECISION
0.268
15 of 56 tested fires improved above noise
HARM RATE
11/56
patches that made the program worse beyond noise
AGENT UTILITY (paired, median)
+1.4%
of the weak→best gap, guided minus no-rubric control
(5–4 states)
Pair-annotation punishes rubrics for real-but-shared/secondary defects;
the causal standard tests each fire by patch-and-rerun. Controls (vague criticism,
domain-swapped rubrics) sit at zero under both.
Caveats that stay attached: 112 of 231 patch attempts
failed to run (selection bias in an unknown direction); eval side is 14 of 60
competitions so far; assist pairs n=10. Library: 183
rubrics mined from all 15 train competitions of split rubric-gen-full75-v1.