← back to the main dashboard · 9/07 updates
On-policy real-bug mining (SWE-Gym → rubric library), per rollout, from the logs of the 241-rollout realgym run · prices from litellm: Opus 5 $5 / $25 per M tokens in / out, Haiku 4.5 $1 / $5, GKE Spot ≈ $0.045 per node-hour
| Component | Basis | Current | Cheap |
|---|---|---|---|
| Haiku agent | 96 LM calls avg × $0.0042; cheap = capped at 150 steps (−20%) | $0.40 | $0.32 |
| Opus judge grading the rollout | ~15 grades × $0.116; same step cap | $1.74 | $1.39 |
| Real test run + pod | Spot e2-standard-4, ~0.5 h | $0.02 | $0.02 |
| Opus distillation | one call per failure; 90% of rollouts fail | $0.09 | $0.09 |
| Embedding + dedup | text-embedding-3-small | <$0.01 | <$0.01 |
| Per rollout | $2.25 | $1.82 | |
| Per rubric | 0.79 rubrics per rollout (241 → 217 failures → 191 after dedup) | $2.85 | $2.30 |
| Precision filter, per candidate | judge vs 454 labelled patches; cheap = Haiku pre-pass, Opus only where it applies, 64 rubrics per call, skip pairs retrieval would never surface | $1.26 | $0.31 |
| Per kept rubric | 78 of 191 survive the filter | $10.0 | $6.4 |
| …if also mined under rc-only feedback | drops the judge from the mining rollout; changes which failures are mined — untested | — | $2.1 |
Where the money is: the Opus judge grading the mining agent's probes is 77% of a rollout. The step cap and the cheaper filter keep the recipe identical; re-using the cached rollouts makes any re-distillation or re-filter ≈ $25 instead of a new $540 run.
For scale: the original manifest-mined library came to ≈ $2 per rubric with 85–90% of curation calls yielding duplicates; a task-specific criterion generated at eval time costs $0.03 and is never reused.