← back to the main dashboard · 9/07 updates

What one rubric costs

On-policy real-bug mining (SWE-Gym → rubric library), per rollout, from the logs of the 241-rollout realgym run · prices from litellm: Opus 5 $5 / $25 per M tokens in / out, Haiku 4.5 $1 / $5, GKE Spot ≈ $0.045 per node-hour

ComponentBasisCurrentCheap
Haiku agent96 LM calls avg × $0.0042; cheap = capped at 150 steps (−20%)$0.40$0.32
Opus judge grading the rollout~15 grades × $0.116; same step cap$1.74$1.39
Real test run + podSpot e2-standard-4, ~0.5 h$0.02$0.02
Opus distillationone call per failure; 90% of rollouts fail$0.09$0.09
Embedding + deduptext-embedding-3-small<$0.01<$0.01
Per rollout$2.25$1.82
Per rubric0.79 rubrics per rollout (241 → 217 failures → 191 after dedup)$2.85$2.30
Precision filter, per candidatejudge vs 454 labelled patches; cheap = Haiku pre-pass, Opus only where it applies, 64 rubrics per call, skip pairs retrieval would never surface$1.26$0.31
Per kept rubric78 of 191 survive the filter$10.0$6.4
…if also mined under rc-only feedbackdrops the judge from the mining rollout; changes which failures are mined — untested—$2.1

Where the money is: the Opus judge grading the mining agent's probes is 77% of a rollout. The step cap and the cheaper filter keep the recipe identical; re-using the cached rollouts makes any re-distillation or re-filter ≈ $25 instead of a new $540 run.

For scale: the original manifest-mined library came to ≈ $2 per rubric with 85–90% of curation calls yielding duplicates; a task-specific criterion generated at eval time costs $0.03 and is never reused.