A measured diagnosis of the CodeWorld GYM framework on KernelBench, end to end.
Across both configurations, every world-model arm loses to plain self-repair. Config A: self-repair 5/20, every WM arm 2/20. Config B: self-repair 6/20, WM arms 2-4/20. The gap is small in absolute terms (n=20, and repeat runs of the same arm differ by ~24 points) but it is consistent in direction across both configs and every WM variant we built.
The race uses 20 real executions where self-repair uses 150, and runs in 107s per task against 764s. That saving is real and reproducible. The third panel is the catch: measured against base — which runs no loop at all — the world-model loop adds zero solves, while the interpreter loop adds three. A 7.5× cost reduction only counts as a win if something is bought with it.
394 world-model-vs-execution audits from live runs. It said FAIL 386
times and PASS 8 times, and it was correct about working code exactly
zero times. False-alarm rate on code that really works: 100%. Balanced
accuracy 48.9% — below chance. A constant function carries no
information, so no threshold, prompt, or rubric database downstream of it
could have worked.
The trap: raw agreement is 91.4%, which
looks excellent. It is meaningless. 93.4% of candidates really fail, so
always answering FAIL scores 93.4%. Every agreement-based check we had was
measuring the base rate.
The world model answers in 3.6s; the real execution takes 110s. With a 31× speed gap the world model wins 100.0% of races (41/41 in config A, 45/45 in B). Every real execution is killed at ~3.6s, so the interpreter never supplies feedback even once. The race arm is therefore behaviourally identical to the WM-only arm, which is exactly what the scores show (2/20 vs 2/20). The one signal that converts tasks is discarded before it is ever read.
On KernelBench the metric is compiled AND correct AND faster than
PyTorch. Of 1476 graded candidates, 34% compile and produce
numerically correct output and simply run too slow — that is 39%
of all failures. Our world model, rubrics and simulator alike, reasons
about correctness. Predicting runtime needs memory bandwidth, occupancy and
tiling, which nothing in the stack models.
From the traces, three of
four sampled predictions were correctness failures on code that was
compiling and correct:
predicted "nvcc build failure" → actually
compiled ✓ correct ✓ speedup 0.12
predicted "out-of-bounds global read" → actually
compiled ✓ correct ✓ speedup 0.80
The real grader was run every turn as a log-only shadow, never fed back.
94-100% of turns change the true score by exactly nothing; mean delta per
turn is +0.0000. The interpreter loop converts 3/20 tasks from fail to
pass. The world-model loop converts 0.
This also falsified
our leading hypothesis. Solve rate correlated +0.861 with turns used, so we
removed the world model's ability to end an episode. Turns went 2.2 →
4.7 and the solve rate stayed at 2/20. The correlation was confounded: the
arm with the most turns was also the arm with the best information.
Splitting criteria by what they measure: functional criteria
(correctness, limits, shapes) have mean lift +0.088; hygiene criteria
(tests, docs, naming, dead code) have mean lift −0.173 —
they flag working code more than broken code. Roughly a third of
criteria point each way and they cancel to a mean lift of −0.028.
The reason is structural: we mined rubrics from pull-request review,
which is a code quality signal, and used them to predict code
correctness. Working code is pragmatic and messy; broken code is
often tidy.
Encouragingly, a readout with per-criterion weights
fitted on held-out data reaches 69.8% balanced accuracy — so the
reward vector does carry signal that uniform "any zero = broken"
aggregation destroys.
Twenty world-model configurations across two model pairings — every
rubric database (general 31k, domain, error-targeted, error+reasons,
online-evolved, generated-frozen), the simulator lane and its debias and
keep-best variants, Opus-5 as the judge, the verdict lane, and the
no-stop change. Dashed lines are the interpreter-only baseline each has to
beat.
Almost everything sits between 0 and 4 of 20 while the
interpreter sits at 5 and 6. The single exception is the last thing we
tried: the evidence-bar debias on the rubric lane, at 5/20 (config A)
and 7/20 (config B) — matching the interpreter on A and passing it
on B.
The shipped prompt flags 73.8% of criteria on code that
demonstrably works, and the debias we already had barely moved it
(56.3%). Adding an evidence bar — a 0 is only permitted if the
model can quote the offending line verbatim, and "I cannot quote it"
resolves to 1 — cuts that to 11.7%.
Scored on both
classes so it is not just a constant-PASS predictor: it passes 61.1% of
working code but only 32.3% of broken code, taking balanced accuracy from
exactly chance (50.0%) to 64.4%. That is the first time this channel
has beaten chance. It has not yet been tested end-to-end on a benchmark.
Each link is measured, not inferred. The efficiency claim survives all of it: the race genuinely costs 7.5× fewer executions and 7× less wall-clock. It just does not currently buy anything.
Benchmark choice is part of the story. We chose KernelBench partly because slow execution makes the race look good — and that same slowness comes from a performance metric the world model structurally cannot predict. On LiveCodeBench, where the metric is "do the tests pass", the framework did win (45.1 vs 33.7 base). That is the question this world model is built to answer.
Statistical power is thin. n=20 per arm, and repeat runs of the same arm differ by ~24 points. The per-turn and per-audit analyses (hundreds to thousands of observations) are far better powered than the solve-rate comparisons, which is why the diagnosis leans on them.