Fresh-era rubric quality across database scaling

One clean run · final split ratio_40_60_final · generated 2026-08-18
Experiment setting
Error precisionError recallPerf precisionPerf recall
0.10.20.30.40.5 Error precision 0.000Error recall 0.000Perf precision 0.234Perf recall 0.160Error precision 0.000Error recall 0.000Perf precision 0.237Perf recall 0.319Error precision 0.152Error recall 0.070Perf precision 0.293Perf recall 0.454Error precision 0.303Error recall 0.240Perf precision 0.276Perf recall 0.427Error precision 0.384Error recall 0.483Perf precision 0.259Perf recall 0.406 0.000.000.230.160.000.000.240.320.150.070.290.450.300.240.280.430.380.480.260.41 stage 5/838.5% / 61.5%74 err · 77 perf rubricsstage 10/1540.0% / 60.0%137 err · 183 perf rubricsstage 15/2339.5% / 60.5%188 err · 259 perf rubricsstage 20/3139.2% / 60.8%278 err · 319 perf rubricsstage 24/3739.3% / 60.7%381 err · 327 perf rubrics
stagemine/eval %err rubricsperf rubricserror Perror Rperf Pperf Rcatchable nperf positives
5/838.5% / 61.5%74770.0000.0000.2340.16018156
10/1540.0% / 60.0%1371830.0000.0000.2370.31958207
15/2339.5% / 60.5%1882590.1520.0700.2930.45486438
20/3139.2% / 60.8%2783190.3030.2400.2760.427121520
24/3739.3% / 60.7%3813270.3840.4830.2590.406149561
Error R = per-program catch rate (≥1 correct-class fire) over catchable failures. Full-scale CIs (10k bootstrap over trajectories): error P [0.31–0.46] · perf P [0.19–0.33] · perf R [0.34–0.49].