New agent architecture · SWE-bench scaling · MLE-bench scaling
1We changed to this agent architecture
The architecture is fully modularized: the agent loop, the model wrapper, the execution
environment and the submission mode are independent pieces, so any one swaps without touching
the others. Execution environments are interchangeable — Modal cloud sandboxes or Docker
containers — and the world model is just another environment: commands still execute in
the sandbox for their side effects, but the observation returned to the agent is the
CWM's rubric grading of the current git diff instead of the interpreter
output. Swapping interpreter ↔ world model changes nothing else.