← main results | causal dashboard P/R comparison + traces sanity + distributions rubric library (183)

Conceptual rubrics — execution-grounded (causal) results

standard: does fixing what the rubric flagged improve the held-out program, decided by the official grader · noise floor 1.54% · generated 2026-08-08 20:55:25 UTC
CAUSAL PRECISION
0.268
15 of 56 tested fires improved above noise
HARM RATE
11/56
patches that made the program worse beyond noise
AGENT UTILITY (paired, median)
+1.4%
of the weak→best gap, guided minus no-rubric control (5–4 states)
Pair-annotation punishes rubrics for real-but-shared/secondary defects; the causal standard tests each fire by patch-and-rerun. Controls (vague criticism, domain-swapped rubrics) sit at zero under both.
Caveats that stay attached: 112 of 231 patch attempts failed to run (selection bias in an unknown direction); eval side is 14 of 60 competitions so far; assist pairs n=10. Library: 183 rubrics mined from all 15 train competitions of split rubric-gen-full75-v1.