← back to the main dashboard

8/31 updates

New agent architecture · SWE-bench scaling · MLE-bench scaling

1We changed to this agent architecture

The architecture is fully modularized: the agent loop, the model wrapper, the execution environment and the submission mode are independent pieces, so any one swaps without touching the others. Execution environments are interchangeable — Modal cloud sandboxes or Docker containers — and the world model is just another environment: commands still execute in the sandbox for their side effects, but the observation returned to the agent is the CWM's rubric grading of the current git diff instead of the interpreter output. Swapping interpreter ↔ world model changes nothing else.

Task — any competition description + official image ① Coding agent (LLM) one bash tool · linear message history limits: 250 steps · $3 · 3 format errors ② Sandbox (Modal / Docker) fresh subshell per command in the task workdir command ALWAYS runs — side effects persist bash cmd the observation channel — an env-only swap, one arm or the other interpreter feedback (baseline) returncode + real output (10k-char cap) ③ CWM feedback (world model) git diff of the change set → top-k rubrics (embedding cosine over the mined rubric library) → an LLM grader marks each 0/1 without executing agent sees violated criteria + predicted reward; interpreter output hidden, returncode kept observation → next turn loop done loop done ④ submit — passes through the swap untouched explicit sentinel command, or the artifact (patch / submission file) collected at stop final artifact → the benchmark's official grader

2We scaled to SWE-bench

What the SWE rubric libraries actually predict

3SWE-bench scaling

SWE-bench Verified resolve rate vs rubric library size

4MLE-bench scaling

MLE-bench Lite medal rate vs rubric library size