Evidence: how the world-model stack fails

Every claim in the diagnosis, with verbatim excerpts pulled from the run logs. Left is what the world model said; right is what actually happened. Nothing is paraphrased. Full traces: traces.html · rubric browser: rubrics.html?v=1785361955

In flight: every rubric run below used the original, maximally biased prompt — all seven of our previous debias runs were --sim-only, so the rubric lane had never actually been debiased. That run (--wm-evidence) is executing now; these traces show the state it is being compared against, not a verdict on debiasing.

The world model condemns code that then passes

394 live audits: it said FAIL on 386 of them and was right about working code 0 times. False-alarm rate on genuinely working code: 100%. Below are real cases where it flagged criteria and the interpreter then passed that exact program.

mlebench/denoising-dirty-documents · mlelite_A_dw
world model said
world model flagged 6 of 10 criteria as VIOLATED

(no reasons logged)
what really happened
the real interpreter then PASSED this exact code (1/1, 2.598s)
mlebench/new-york-city-taxi-fare-prediction · mlelite_A_dw
world model said
world model flagged 9 of 10 criteria as VIOLATED

(no reasons logged)
what really happened
the real interpreter then PASSED this exact code (1/1, 2.229s)
mlebench/ranzcr-clip-catheter-line-classification · mlelite_A_dw
world model said
world model flagged 10 of 10 criteria as VIOLATED

(no reasons logged)
what really happened
the real interpreter then PASSED this exact code (1/1, 8.96s)
mlebench/denoising-dirty-documents · mlelite_A_errR
world model said
world model flagged 5 of 10 criteria as VIOLATED

**The task is not from the benchmark family the rubric targets.** The target error types describe failures of a *coding agent solv
      -> WM: Task is a Kaggle OCR denoising competition, which matches the data-science benchmark family the rubric targets.

The solution loads and parses the actual competition data files before emitting any submission, and never writes a constant, empty
      -> WM: Code writes fallback empty/sample-copied submission when data loading fails (lines 330-341, 354-359, 368-373, 385-390, 408-413).

**The solutions contain none of the mechanisms the error types describe.** Both ranked snippets differ only in whether they guard 
      -> WM: Rubric describes non-existent Keras code; this is a Kaggle OCR task, unrelated to the stated criterion.
what really happened
the real interpreter then PASSED this exact code (1/1, 1.655s)
mlebench/mlsp-2013-birds · mlelite_A_errR
world model said
world model flagged 8 of 10 criteria as VIOLATED

The solution verifies that any generated submission is non-empty and populated with real predictions (correct row count and value 
      -> WM: Code ends abruptly at line 289 in main(); submission verification logic is incomplete/missing.

**The task is not from the benchmark family the rubric targets.** The target error types describe failures of a *coding agent solv
      -> WM: Task context describes Kaggle bird classification (multi-label audio), not Keras tensor resizing; mismatch between rubric and task.

Discard everything else "however good it sounds," there are **no valid observations to keep**. The genuine issues one could raise 
      -> WM: Rubric designed for Keras ops bug-fix; this is a Kaggle competition solution. No valid observations map to listed error types.
what really happened
the real interpreter then PASSED this exact code (1/1, 6.632s)
mlebench/ranzcr-clip-catheter-line-classification · mlelite_A_errR
world model said
world model flagged 10 of 10 criteria as VIOLATED

**The solutions contain none of the mechanisms the error types describe.** Both ranked snippets differ only in whether they guard 
      -> WM: Rubric does not apply to this Kaggle competition task. Code submission/data handling, not Keras library fix.

**The task is not from the benchmark family the rubric targets.** The target error types describe failures of a *coding agent solv
      -> WM: Rubric targets Keras library bug-fix; this is Kaggle competition code. Different task domains entirely.

Discard everything else "however good it sounds," there are **no valid observations to keep**. The genuine issues one could raise 
      -> WM: Rubric and code task mismatch. No valid error-type mappings exist between them.
what really happened
the real interpreter then PASSED this exact code (1/1, 3.345s)
mlebench/tabular-playground-series-dec-2021 · mlelite_A_errR
world model said
world model flagged 3 of 10 criteria as VIOLATED

**The task is not from the benchmark family the rubric targets.** The target error types describe failures of a *coding agent solv
      -> WM: Rubric targets Keras ops.image.resize library bug; code is Kaggle competition solver. Task mismatch.

**The solutions contain none of the mechanisms the error types describe.** Both ranked snippets differ only in whether they guard 
      -> WM: Rubric error types (symbolic tensors, type guards) irrelevant to this competition-solving code.

The solution loads and parses the actual competition data files before emitting any submission, and never writes a constant, empty
      -> WM: Line 139–142: fallback creates empty/sample-copied submission when train.csv or test.csv missing.
what really happened
the real interpreter then PASSED this exact code (1/1, 4.922s)
mlebench/aptos2019-blindness-detection · mlelite_A_err
world model said
world model flagged 2 of 10 criteria as VIOLATED

(no reasons logged)
what really happened
the real interpreter then PASSED this exact code (1/1, 7.227s)

It predicts the wrong KIND of failure

34% of KernelBench candidates compile and are numerically correct and merely run too slow — 39% of all failures. The world model reasons about correctness, so it invents compile errors and out-of-bounds reads that are not there.

kernelbench/l2_12_Gemm_Multiply_LeakyReLU · kernelbench_A_fixAB
world model said
WM PREDICTED:
Runtime error in forward(): The fused_gemm_bias_scale_leakyrelu_cuda kernel will be called with uninitialized weight and bias tensors. This will produce incorrect numerical results (undefined behavior from uninitialized memory). Additionally, there is a missing initialization step - the weight and bias buffers should be initialized with values from nn.Linear, but they are only registered as empty buffers. The nn.Linear layer that would normally initialize these is absent from __init__.
what really happened
INTERPRETER FOUND:
kernel eval compiled + correct + speedup>1.0 correct but speedup=0.01 <= 1.0 (kernel 272.0 ms vs ref 2.72 ms) {'compiled': True, 'correctness': True, 'runtime_ms': 272.0, 'ref_runtime_ms': 2.72, 'speedup': 0.01, 'excessive_speedup': False, 'static_check': {'valid': True, 'errors': [], 'warnings': []}, 'metadata': {'hardware': 'NVIDIA H200', 'device': '0', 'correctness_trials': '(5 / 5)'}, 'error': None}
kernelbench/l2_17_Conv2d_InstanceNorm_Divide · kernelbench_A_fixAB
world model said
WM PREDICTED:
Runtime error in InstanceNorm2dCustom.forward() at the line: `output = instance_norm_ops.instance_norm_forward_cuda(input, mean_expanded, var_expanded, self.eps, divide_by)`. The custom CUDA kernel instance_norm_forward_kernel expects mean and var tensors with shape (batch_size * channels,) but the code passes mean_expanded and var_expanded with shape (batch_size, channels, 1, 1). This shape mismatch will cause the kernel to write to incorrect memory locations or access out-of-bounds memory, resulting in a CUDA err
what really happened
INTERPRETER FOUND:
kernel eval compiled + correct + speedup>1.0 correct but speedup=0.4221 <= 1.0 (kernel 10.4 ms vs ref 4.39 ms) {'compiled': True, 'correctness': True, 'runtime_ms': 10.4, 'ref_runtime_ms': 4.39, 'speedup': 0.4221, 'excessive_speedup': False, 'static_check': {'valid': True, 'errors': [], 'warnings': ['Uses torch.nn compute layer (only containers, Parameter, init allowed)']}, 'metadata': {'hardware': 'NVIDIA H200', 'de
kernelbench/l1_10_3D_tensor_matrix_multiplication · kernelbench_A_fixB2
world model said
WM PREDICTED:
RUNTIME LOGIC ERROR: The kernel has a critical indexing bug in the B_tile loading logic. When loading B_tile[ty][tx], the code uses (t * TILE_K + ty) as the K-dimension index and tx as the L-dimension index. However, tx ranges from 0 to TILE_L-1 (0 to 31), but B's second dimension is L which is 768. This means the thread block is dimensioned as (TILE_L=32, TILE_M=32) = (32, 32) threads = 1024 threads, but the B_tile loading assumes tx maps to the column index of B. The actual bug is more subtle: the thread block is
what really happened
INTERPRETER FOUND:
kernel eval compiled + correct + speedup>1.0 correct but speedup=0.1079 <= 1.0 (kernel 9.64 ms vs ref 1.04 ms) {'compiled': True, 'correctness': True, 'runtime_ms': 9.64, 'ref_runtime_ms': 1.04, 'speedup': 0.1079, 'excessive_speedup': False, 'static_check': {'valid': True, 'errors': [], 'warnings': []}, 'metadata': {'hardware': 'NVIDIA H200', 'device': '0', 'correctness_trials': '(5 / 5)'}, 'error': None}
kernelbench/l1_13_Matmul_for_symmetric_matrices · kernelbench_A_fixB2
world model said
WM PREDICTED:
CORRECTNESS ISSUE: The custom kernel computes C = A * B^T (where B is treated as if transposed in the computation loop), not C = A * B. In the symmetric_matmul_upper_kernel, the line `sum += A[row * N + k] * B[col * N + k];` treats B as if indexing by [col][k], but this computes A[row][k] * B[col][k] which is equivalent to A[row][k] * B[k][col] due to the symmetry comment. However, the correct matrix multiplication should be `sum += A[row * N + k] * B[k * N + col];` to compute C[row][col] = sum_k A[row][k] * B[k][c
what really happened
INTERPRETER FOUND:
kernel eval compiled + correct + speedup>1.0 correct but speedup=0.1718 <= 1.0 (kernel 15.6 ms vs ref 2.68 ms) {'compiled': True, 'correctness': True, 'runtime_ms': 15.6, 'ref_runtime_ms': 2.68, 'speedup': 0.1718, 'excessive_speedup': False, 'static_check': {'valid': True, 'errors': [], 'warnings': []}, 'metadata': {'hardware': 'NVIDIA H200', 'device': '0', 'correctness_trials': '(5 / 5)'}, 'error': None}
kernelbench/l1_17_Matmul_with_transposed_B · kernelbench_A_fixB2
world model said
WM PREDICTED:
CUDA runtime error or incorrect numerical results due to semantic mismatch: The kernel expects B to be (K, N) but receives B as (N, K). The memory indexing Bs[ty][tx] = B[(k + ty) * N + col] is incorrect because it assumes B has N columns when B actually has K columns. This causes out-of-bounds or incorrect memory access. The fix would require either: (1) transposing B before passing to kernel, or (2) rewriting kernel to read B[(k + ty) * K + col] or adjusting the entire computation logic.
what really happened
INTERPRETER FOUND:
kernel eval compiled + correct + speedup>1.0 correct but speedup=0.1107 <= 1.0 (kernel 24.3 ms vs ref 2.69 ms) {'compiled': True, 'correctness': True, 'runtime_ms': 24.3, 'ref_runtime_ms': 2.69, 'speedup': 0.1107, 'excessive_speedup': False, 'static_check': {'valid': True, 'errors': [], 'warnings': []}, 'metadata': {'hardware': 'NVIDIA H200', 'device': '0', 'correctness_trials': '(5 / 5)'}, 'error': None}
kernelbench/l1_18_Matmul_with_transposed_both · kernelbench_A_fixB2
world model said
WM PREDICTED:
CRITICAL RUNTIME ERROR: The code has a fundamental mathematical error. The original forward method computes torch.matmul(A.T, B.T), but the candidate code implements torch.matmul(A, B) instead. This is a semantic mismatch that produces incorrect results. Additionally, there are implementation bugs: (1) The CUDA code calls B_t.t() and A_t.t() which perform transposes on already-transposed tensors, creating double transposes - this is inefficient and may cause incorrect leading dimensions. (2) The cuBLAS call uses in
what really happened
INTERPRETER FOUND:
kernel eval compiled + correct + speedup>1.0 correct but speedup=0.1053 <= 1.0 (kernel 26.3 ms vs ref 2.77 ms) {'compiled': True, 'correctness': True, 'runtime_ms': 26.3, 'ref_runtime_ms': 2.77, 'speedup': 0.1053, 'excessive_speedup': False, 'static_check': {'valid': True, 'errors': [], 'warnings': []}, 'metadata': {'hardware': 'NVIDIA H200', 'device': '0', 'correctness_trials': '(5 / 5)'}, 'error': None}
kernelbench/l2_18_Matmul_Sum_Max_AvgPool_LogSumExp_LogSumExp · kernelbench_A_fixB2
world model said
WM PREDICTED:
1. Line 'import math' appears AFTER class definition (at end of file) instead of at the beginning. This will cause NameError when ModelNew.reset_parameters() tries to use math.sqrt(5) during __init__.

2. In the CUDA kernel fused_model_kernel, line 'float* local_out = shared_mem;' allocates shared_mem but the kernel is launched with shared_mem_size = 0 (last argument to <<<>>>). This declares shared_mem as 'extern __shared__' but zero bytes are allocated, causing undefined behavior or segmentation fault when access
what really happened
INTERPRETER FOUND:
kernel eval compiled + correct + speedup>1.0 correct but speedup=0.0097 <= 1.0 (kernel 285.0 ms vs ref 2.76 ms) {'compiled': True, 'correctness': True, 'runtime_ms': 285.0, 'ref_runtime_ms': 2.76, 'speedup': 0.0097, 'excessive_speedup': False, 'static_check': {'valid': True, 'errors': [], 'warnings': []}, 'metadata': {'hardware': 'NVIDIA H200', 'device': '0', 'correctness_trials': '(5 / 5)'}, 'error': None}
kernelbench/l1_11_4D_tensor_matrix_multiplication · kernelbench_A_fixB
world model said
WM PREDICTED:
CRITICAL KERNEL BUG: The kernel accesses B with incorrect indexing. Line 'sum += A[row * K_in + l] * B[l * N + col];' assumes B is in row-major format with stride N, but B is actually in row-major format with stride k (the second dimension). Since B has shape (l, k) = (256, 768), B[l, k] in row-major is stored as B_flat[l*768 + k], not B_flat[l*N + col]. When col ranges from 0-767 and N=768, this happens to work correctly in this specific case, BUT the kernel logic is semantically incorrect and uses the global N va
what really happened
INTERPRETER FOUND:
kernel eval compiled + correct + speedup>1.0 correct but speedup=0.1758 <= 1.0 (kernel 62.0 ms vs ref 10.9 ms) {'compiled': True, 'correctness': True, 'runtime_ms': 62.0, 'ref_runtime_ms': 10.9, 'speedup': 0.1758, 'excessive_speedup': False, 'static_check': {'valid': True, 'errors': [], 'warnings': []}, 'metadata': {'hardware': 'NVIDIA H200', 'device': '0', 'correctness_trials': '(5 / 5)'}, 'error': None}

The race is never actually a race

The world model answers in 3.6s; a real execution takes 110s. It wins 100.0% of races (41/41 in config A, 45/45 in B), so every real execution is killed a few seconds in and the interpreter never speaks.

kernelbench/l1_100_HingeLoss · kernelbench_A_tr
world model said
races this episode: turn 0 -> world_model (3.446s), turn 1 -> world_model (4.101s), turn 2 -> world_model (4.075s), turn 3 -> world_model (3.323s)
what really happened
real executions KILLED mid-flight: turn 0 after 3.446s, turn 1 after 4.1s, turn 2 after 4.075s, turn 3 after 3.322s
(a full execution takes ~110s, so none got near finishing)
kernelbench/l1_10_3D_tensor_matrix_multiplication · kernelbench_A_tr
world model said
races this episode: turn 0 -> world_model (3.643s), turn 1 -> world_model (3.74s)
what really happened
real executions KILLED mid-flight: turn 0 after 3.642s, turn 1 after 3.739s
(a full execution takes ~110s, so none got near finishing)
kernelbench/l1_11_4D_tensor_matrix_multiplication · kernelbench_A_tr
world model said
races this episode: turn 0 -> world_model (3.559s)
what really happened
real executions KILLED mid-flight: turn 0 after 3.559s
(a full execution takes ~110s, so none got near finishing)
kernelbench/l1_12_Matmul_with_diagonal_matrices_ · kernelbench_A_tr
world model said
races this episode: turn 0 -> world_model (3.239s), turn 1 -> world_model (3.203s)
what really happened
real executions KILLED mid-flight: turn 0 after 3.237s, turn 1 after 3.202s
(a full execution takes ~110s, so none got near finishing)
kernelbench/l1_13_Matmul_for_symmetric_matrices · kernelbench_A_tr
world model said
races this episode: turn 0 -> world_model (3.741s), turn 1 -> world_model (3.732s), turn 2 -> world_model (3.157s)
what really happened
real executions KILLED mid-flight: turn 0 after 3.74s, turn 1 after 3.731s, turn 2 after 3.155s
(a full execution takes ~110s, so none got near finishing)
kernelbench/l1_14_Matmul_for_upper_triangular_matrices · kernelbench_A_tr
world model said
races this episode: turn 0 -> world_model (3.704s), turn 1 -> world_model (3.373s)
what really happened
real executions KILLED mid-flight: turn 0 after 3.703s, turn 1 after 3.373s
(a full execution takes ~110s, so none got near finishing)
kernelbench/l1_15_Matmul_for_lower_triangular_matrices · kernelbench_A_tr
world model said
races this episode: turn 0 -> world_model (3.15s), turn 1 -> world_model (3.489s)
what really happened
real executions KILLED mid-flight: turn 0 after 3.149s, turn 1 after 3.488s
(a full execution takes ~110s, so none got near finishing)
kernelbench/l1_16_Matmul_with_transposed_A · kernelbench_A_tr
world model said
races this episode: turn 0 -> world_model (3.378s), turn 1 -> world_model (3.476s)
what really happened
real executions KILLED mid-flight: turn 0 after 3.377s, turn 1 after 3.475s
(a full execution takes ~110s, so none got near finishing)

Turn after turn of feedback moves the true score by nothing

The real grader ran every turn as a log-only shadow. 94–100% of turns change the true score by exactly zero; mean delta +0.0000. These episodes each spent several turns rewriting code with no change in grade.

kernelbench/l1_100_HingeLoss · kernelbench_A_dndw
world model said
turn 0 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=0] The solution avoids introducing unnecessary computational overhead per invocation, favoring vectorized or kernel-native operations over manual element

turn 1 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=0] The solution avoids introducing unnecessary computational overhead per invocation, favoring vectorized or kernel-native operations over manual element

turn 2 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will attempt to compile and load the CUDA kernels. If compilation succeeds, the ModelNew class will be ins
what really happened
true score after every one of those 3 turns: 0.00, 0.00, 0.00

the code was re-written repeatedly and the real grade never changed
kernelbench/l1_10_3D_tensor_matrix_multiplication · kernelbench_A_dndw
world model said
turn 0 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will compile successfully and execute the batched matrix multiplication. For inputs A of shape (16, 1024, 

turn 1 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will compile successfully and execute the custom CUDA kernel for batched matrix multiplication. For inputs

turn 2 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=1] The solution provides the compute kernel(s) (e.g., dot-product) needed to operate on the new format for the targeted execution paths.
  [reward=0] The s
what really happened
true score after every one of those 6 turns: 0.00, 0.00, 0.00, 0.00, 0.00, 0.00

the code was re-written repeatedly and the real grade never changed
kernelbench/l1_11_4D_tensor_matrix_multiplication · kernelbench_A_dndw
world model said
turn 0 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=1] The solution provides the compute kernel(s) (e.g., dot-product) needed to operate on the new format for the targeted execution paths.
  [reward=0] The s

turn 1 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=0] The solution avoids introducing unnecessary computational overhead per invocation, favoring vectorized or kernel-native operations over manual element

turn 2 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will compile successfully and execute without runtime errors. The ModelNew class will perform 4D tensor-ma
what really happened
true score after every one of those 6 turns: 0.00, 0.00, 0.00, 0.00, 0.00, 0.00

the code was re-written repeatedly and the real grade never changed
kernelbench/l1_11_4D_tensor_matrix_multiplication · kernelbench_A_dndw
world model said
turn 0 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=1] The solution provides the compute kernel(s) (e.g., dot-product) needed to operate on the new format for the targeted execution paths.
  [reward=0] The s

turn 1 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will compile successfully and execute, producing output tensor of shape (b, i, j, k) = (8, 256, 512, 768) 

turn 2 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=0] The solution provides the compute kernel(s) (e.g., dot-product) needed to operate on the new format for the targeted execution paths.
  [reward=1] The s
what really happened
true score after every one of those 4 turns: 0.00, 0.00, 0.00, 0.00

the code was re-written repeatedly and the real grade never changed
kernelbench/l1_13_Matmul_for_symmetric_matrices · kernelbench_A_dndw
world model said
turn 0 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=0] The solution avoids introducing unnecessary computational overhead per invocation, favoring vectorized or kernel-native operations over manual element

turn 1 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=1] The solution avoids introducing unnecessary computational overhead per invocation, favoring vectorized or kernel-native operations over manual element

turn 2 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=1] The solution avoids introducing unnecessary computational overhead per invocation, favoring vectorized or kernel-native operations over manual element
what really happened
true score after every one of those 8 turns: 0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00

the code was re-written repeatedly and the real grade never changed
kernelbench/l1_16_Matmul_with_transposed_A · kernelbench_A_dndw
world model said
turn 0 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=1] The solution avoids introducing unnecessary computational overhead per invocation, favoring vectorized or kernel-native operations over manual element

turn 1 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=1] The solution avoids introducing unnecessary computational overhead per invocation, favoring vectorized or kernel-native operations over manual element

turn 2 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will compile successfully and execute without runtime errors. The ModelNew class will correctly compute C 
what really happened
true score after every one of those 5 turns: 0.00, 0.00, 0.00, 0.00, 0.00

the code was re-written repeatedly and the real grade never changed
kernelbench/l1_16_Matmul_with_transposed_A · kernelbench_A_dndw
world model said
turn 0 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: Code will compile successfully and execute without runtime errors. The custom CUDA kernel will produce mathematical

turn 1 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will compile successfully and execute without runtime errors. It will produce a tensor of shape (M, N) = (

turn 2 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will compile successfully and produce a ModelNew class that can be instantiated and called with input tens
what really happened
true score after every one of those 4 turns: 0.00, 0.00, 0.00, 0.00

the code was re-written repeatedly and the real grade never changed
kernelbench/l1_17_Matmul_with_transposed_B · kernelbench_A_dndw
world model said
turn 0 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will compile successfully and execute without runtime errors. The custom CUDA kernel will compute matrix m

turn 1 feedback:
  The code world model SIMULATED running your previous code (prediction, not a real run):
  Predicted output/behavior: The code will compile successfully and execute the matrix multiplication operation. For inputs A (2048, 8192) and B

turn 2 feedback:
  Rubric rewards for your previous code (1 = satisfied, 0 = violated):
  [reward=1] The solution provides the compute kernel(s) (e.g., dot-product) needed to operate on the new format for the targeted execution paths.
  [reward=0] The s
what really happened
true score after every one of those 6 turns: 0.00, 0.00, 0.00, 0.00, 0.00, 0.00

the code was re-written repeatedly and the real grade never changed

A third of the rubric criteria point the wrong way

Hygiene criteria (tests, docs, naming) have mean lift −0.173: they flag working code MORE than broken code, because working code is pragmatic and messy while broken code is often tidy. They cancel the functional criteria (+0.088) to a net −0.028.

criterion 9bbd36551e · all runs pooled
world model said
The solution is structured clearly, with helper routines or well-named components that make the partitioning and scoring logic readable and maintainable.
what really happened
flagged on 4/8 WORKING programs (50%)
flagged on 0/11 BROKEN programs (0%)

lift -0.50 — it points the WRONG WAY
criterion 5395d4a0eb · all runs pooled
world model said
The solution is structured clearly with well-named helper routines separating candidate generation, validation, and output.
what really happened
flagged on 3/5 WORKING programs (60%)
flagged on 2/14 BROKEN programs (14%)

lift -0.46 — it points the WRONG WAY
criterion 39668c10ee · all runs pooled
world model said
Discard everything else "however good it sounds," there are **no valid observations to keep**. The genuine issues one could raise about these solutions (e.g., Rank 2 increments the counter but never repositions the summary writer via `_pop_writer`/`_push_writer`, so scalars may still not advance in the writer's step context) are real code-review points, but none of them correspond to any target error type. Forcing th
what really happened
flagged on 4/5 WORKING programs (80%)
flagged on 5/6 BROKEN programs (83%)

lift +0.03 — it points the WRONG WAY
criterion 6d22564846 · all runs pooled
world model said
**The task is not from the benchmark family the rubric targets.** The target error types describe failures of a *coding agent solving data-science / Kaggle-style competition tasks* (writing submissions, loading training data, importing ML libraries, respecting compute budgets, handling dataset layouts). The actual `<request>` is a small library bug-fix in Keras (`ops.image.resize` allowing dynamic tensor sizes). Ther
what really happened
flagged on 5/6 WORKING programs (83%)
flagged on 8/9 BROKEN programs (89%)

lift +0.06 — it points the WRONG WAY
criterion e14cf3d1dd · all runs pooled
world model said
**The solutions contain none of the mechanisms the error types describe.** Both ranked snippets differ only in whether they guard `size[0]`/`size[1]` comparisons with `isinstance(x, int)` before applying `<= 0` — i.e., allowing symbolic/dynamic (non-int) sizes. This maps to input-type validation, not to any of: - silent-fallback-invalid-submission - unvalidated-column-schema - missing-third-party-dependency - no-comp
what really happened
flagged on 4/5 WORKING programs (80%)
flagged on 7/8 BROKEN programs (88%)

lift +0.07 — it points the WRONG WAY

What the interpreter says, that the world model cannot

The interpreter's feedback names the defect with numbers. This is the signal that converts 3/20 tasks from fail to pass — and the signal the race discards at 3.6s in favour of a prediction.

kernelbench/l1_100_HingeLoss · kernelbench_A_tr
world model said
INTERPRETER feedback (self-repair arm):
Real execution: 0/1 tests passed.
- input: kernel eval
  expected: compiled + correct + speedup>1.0
  got: eval error: eval_kernel_against_ref returned None twice (build lock error)
what really happened
names the concrete defect, with numbers.
this is the feedback the WM arms cancel at 3.6s and replace with a prediction.
kernelbench/l1_10_3D_tensor_matrix_multiplication · kernelbench_A_tr
world model said
INTERPRETER feedback (self-repair arm):
Real execution: 0/1 tests passed.
- input: kernel eval
  expected: compiled + correct + speedup>1.0
  got: eval error: /site-packages/torch/cuda/__init__.py", line 1162, in synchronize
    return torch._C._cuda_synchronize()
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
torch.AcceleratorError: CUDA error: an ill
what really happened
names the concrete defect, with numbers.
this is the feedback the WM arms cancel at 3.6s and replace with a prediction.
kernelbench/l1_11_4D_tensor_matrix_multiplication · kernelbench_A_tr
world model said
INTERPRETER feedback (self-repair arm):
Real execution: 0/1 tests passed.
- input: kernel eval
  expected: compiled + correct + speedup>1.0
  got: eval error: eval_kernel_against_ref returned None twice (build lock error)
what really happened
names the concrete defect, with numbers.
this is the feedback the WM arms cancel at 3.6s and replace with a prediction.
kernelbench/l1_12_Matmul_with_diagonal_matrices_ · kernelbench_A_tr
world model said
INTERPRETER feedback (self-repair arm):
Real execution: 0/1 tests passed.
- input: kernel eval
  expected: compiled + correct + speedup>1.0
  got: correct but speedup=0.8963 <= 1.0 (kernel 0.0627 ms vs ref 0.0562 ms)
what really happened
names the concrete defect, with numbers.
this is the feedback the WM arms cancel at 3.6s and replace with a prediction.
kernelbench/l1_13_Matmul_for_symmetric_matrices · kernelbench_A_tr
world model said
INTERPRETER feedback (self-repair arm):
Real execution: 0/1 tests passed.
- input: kernel eval
  expected: compiled + correct + speedup>1.0
  got: wrong output: Output mismatch (0 / 5)
what really happened
names the concrete defect, with numbers.
this is the feedback the WM arms cancel at 3.6s and replace with a prediction.
kernelbench/l1_14_Matmul_for_upper_triangular_matrices · kernelbench_A_tr
world model said
INTERPRETER feedback (self-repair arm):
Real execution: 0/1 tests passed.
- input: kernel eval
  expected: compiled + correct + speedup>1.0
  got: correct but speedup=0.6267 <= 1.0 (kernel 4.34 ms vs ref 2.72 ms)
what really happened
names the concrete defect, with numbers.
this is the feedback the WM arms cancel at 3.6s and replace with a prediction.
kernelbench/l1_15_Matmul_for_lower_triangular_matrices · kernelbench_A_tr
world model said
INTERPRETER feedback (self-repair arm):
Real execution: 0/1 tests passed.
- input: kernel eval
  expected: compiled + correct + speedup>1.0
  got: correct but speedup=0.3953 <= 1.0 (kernel 6.88 ms vs ref 2.72 ms)
what really happened
names the concrete defect, with numbers.
this is the feedback the WM arms cancel at 3.6s and replace with a prediction.
kernelbench/l1_16_Matmul_with_transposed_A · kernelbench_A_tr
world model said
INTERPRETER feedback (self-repair arm):
Real execution: 0/1 tests passed.
- input: kernel eval
  expected: compiled + correct + speedup>1.0
  got: correct but speedup=0.1019 <= 1.0 (kernel 25.7 ms vs ref 2.62 ms)
what really happened
names the concrete defect, with numbers.
this is the feedback the WM arms cancel at 3.6s and replace with a prediction.