Every rubric the world model can retrieve, as the judge reads it. Slices are nested prefixes (10 ⊂ 100 ⊂ 1000 ⊂ all) in mining order. Each page has a text filter and shows the raw criterion block plus its provenance.
| benchmark | library | rubrics | what it is |
|---|---|---|---|
| SWE-bench | swe-rubrics-mined-1000 | 1000 | 1000 rubrics mined from SWE-smith / SWE-Gym rollouts (state-based recipe); the 10 / 100 slices are its first 10 / 100 records |
| SWE-bench | swe-rubrics-mined-100 | 100 | first 100 of the 1000 |
| SWE-bench | swe-rubrics-mined-10 | 10 | first 10 of the 1000 |
| SWE-bench | swe-rubrics-onpolicy-gym | 77 | 77 rubrics, failure-based recipe; only the pool for the random-12 baseline |
| MLE-bench | mle-rubrics-onpolicy-all | 1076 | 1076 rubrics: one Opus lesson per failed rollout on the DA pool (InfiAgent-DABench + DA-Code), dedup at 0.90 |
| MLE-bench | mle-rubrics-onpolicy-1000 | 1000 | first 1000 of all |
| MLE-bench | mle-rubrics-onpolicy-100 | 100 | first 100 of all |
| MLE-bench | mle-rubrics-onpolicy-10 | 10 | first 10 of all |
| SWE-fficiency | perf-rubrics-onpolicy-all | 69 | 69 rubrics mined from the perf pool (injected slowdowns on SWE-smith repos) |
| SWE-fficiency | perf-rubrics-onpolicy-10 | 10 | first 10 of all |
Record schema and recipes: mining/README.md in the repo. Local copies: data/libraries/*.json.