Why the world-model stack loses to the interpreter

A measured diagnosis of the CodeWorld GYM framework on KernelBench, end to end.

Every number below is computed from run logs by scripts/make_diagnosis_report.py. Full verbatim traces: traces.html

0 / 26times the world model was right about working code
48.9%balanced accuracy — below chance
100%of races won by the WM, so the interpreter never speaks
7.5×fewer real executions — the claim that survives

1. The result we have to explain

Across both configurations, every world-model arm loses to plain self-repair. Config A: self-repair 5/20, every WM arm 2/20. Config B: self-repair 6/20, WM arms 2-4/20. The gap is small in absolute terms (n=20, and repeat runs of the same arm differ by ~24 points) but it is consistent in direction across both configs and every WM variant we built.

1. The result we have to explain

2. What DOES work: the cost side

The race uses 20 real executions where self-repair uses 150, and runs in 107s per task against 764s. That saving is real and reproducible. The third panel is the catch: measured against base — which runs no loop at all — the world-model loop adds zero solves, while the interpreter loop adds three. A 7.5× cost reduction only counts as a win if something is bought with it.

2. What DOES work: the cost side

3. Root cause: the world model is a constant predictor

394 world-model-vs-execution audits from live runs. It said FAIL 386 times and PASS 8 times, and it was correct about working code exactly zero times. False-alarm rate on code that really works: 100%. Balanced accuracy 48.9% — below chance. A constant function carries no information, so no threshold, prompt, or rubric database downstream of it could have worked.

The trap: raw agreement is 91.4%, which looks excellent. It is meaningless. 93.4% of candidates really fail, so always answering FAIL scores 93.4%. Every agreement-based check we had was measuring the base rate.

3. Root cause: the world model is a constant predictor

4. Why the race never actually races

The world model answers in 3.6s; the real execution takes 110s. With a 31× speed gap the world model wins 100.0% of races (41/41 in config A, 45/45 in B). Every real execution is killed at ~3.6s, so the interpreter never supplies feedback even once. The race arm is therefore behaviourally identical to the WM-only arm, which is exactly what the scores show (2/20 vs 2/20). The one signal that converts tasks is discarded before it is ever read.

4. Why the race never actually races

5. The world model is answering the wrong question

On KernelBench the metric is compiled AND correct AND faster than PyTorch. Of 1476 graded candidates, 34% compile and produce numerically correct output and simply run too slow — that is 39% of all failures. Our world model, rubrics and simulator alike, reasons about correctness. Predicting runtime needs memory bandwidth, occupancy and tiling, which nothing in the stack models.

From the traces, three of four sampled predictions were correctness failures on code that was compiling and correct:
predicted "nvcc build failure" → actually compiled ✓ correct ✓ speedup 0.12
predicted "out-of-bounds global read" → actually compiled ✓ correct ✓ speedup 0.80

5. The world model is answering the wrong question

6. The feedback does not move the code

The real grader was run every turn as a log-only shadow, never fed back. 94-100% of turns change the true score by exactly nothing; mean delta per turn is +0.0000. The interpreter loop converts 3/20 tasks from fail to pass. The world-model loop converts 0.

This also falsified our leading hypothesis. Solve rate correlated +0.861 with turns used, so we removed the world model's ability to end an episode. Turns went 2.2 → 4.7 and the solve rate stayed at 2/20. The correlation was confounded: the arm with the most turns was also the arm with the best information.

6. The feedback does not move the code

7. Why the rubric channel specifically carries no signal

Splitting criteria by what they measure: functional criteria (correctness, limits, shapes) have mean lift +0.088; hygiene criteria (tests, docs, naming, dead code) have mean lift −0.173 — they flag working code more than broken code. Roughly a third of criteria point each way and they cancel to a mean lift of −0.028.

The reason is structural: we mined rubrics from pull-request review, which is a code quality signal, and used them to predict code correctness. Working code is pragmatic and messy; broken code is often tidy.

Encouragingly, a readout with per-criterion weights fitted on held-out data reaches 69.8% balanced accuracy — so the reward vector does carry signal that uniform "any zero = broken" aggregation destroys.

7. Why the rubric channel specifically carries no signal

8. Everything we tried, on one axis

Twenty world-model configurations across two model pairings — every rubric database (general 31k, domain, error-targeted, error+reasons, online-evolved, generated-frozen), the simulator lane and its debias and keep-best variants, Opus-5 as the judge, the verdict lane, and the no-stop change. Dashed lines are the interpreter-only baseline each has to beat.

Almost everything sits between 0 and 4 of 20 while the interpreter sits at 5 and 6. The single exception is the last thing we tried: the evidence-bar debias on the rubric lane, at 5/20 (config A) and 7/20 (config B) — matching the interpreter on A and passing it on B.

8. Everything we tried, on one axis

9. The fix behind that result

The shipped prompt flags 73.8% of criteria on code that demonstrably works, and the debias we already had barely moved it (56.3%). Adding an evidence bar — a 0 is only permitted if the model can quote the offending line verbatim, and "I cannot quote it" resolves to 1 — cuts that to 11.7%.

Scored on both classes so it is not just a constant-PASS predictor: it passes 61.1% of working code but only 32.3% of broken code, taking balanced accuracy from exactly chance (50.0%) to 64.4%. That is the first time this channel has beaten chance. It has not yet been tested end-to-end on a benchmark.

9. The fix behind that result

The chain, end to end

  1. The world model is a constant FAIL predictor (0/26 correct on working code, 48.9% balanced accuracy).
  2. So its feedback names a defect that is not there — and on 39% of KernelBench failures it is reasoning about correctness when the real problem is speed, which it cannot model at all.
  3. So the coder edits the wrong thing: 94.5% of turns change the true score by zero.
  4. So the loop converts 0 tasks from fail to pass, against the interpreter's 3.
  5. And because the world model is 31× faster, it wins every race and the interpreter's feedback — the only feedback that converts anything — is cancelled at 3.6s and never read.

Each link is measured, not inferred. The efficiency claim survives all of it: the race genuinely costs 7.5× fewer executions and 7× less wall-clock. It just does not currently buy anything.

What we ruled out along the way

Two honest caveats

Benchmark choice is part of the story. We chose KernelBench partly because slow execution makes the race look good — and that same slowness comes from a performance metric the world model structurally cannot predict. On LiveCodeBench, where the metric is "do the tests pass", the framework did win (45.1 vs 33.7 base). That is the question this world model is built to answer.

Statistical power is thin. n=20 per arm, and repeat runs of the same arm differ by ~24 points. The per-turn and per-audit analyses (hundreds to thousands of observations) are far better powered than the solve-rate comparisons, which is why the diagnosis leans on them.