AI-Assisted Software DevelopmentAug 20, 2026
StateM claims 95.3% on Terminal-Bench 2.1 by scaling the harness, not the model
A paper submitted 15 August 2026 by Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang and Kai Wang argues that long-horizon agents fail on tasks whose individual steps their model can already solve, and attacks the runtime rather than the weights. StateM organises execution around durable states, recoverable runbooks and enforceable procedural controls. Reported figures: GPT-5.6 Sol xhigh at 95.3% raw accuracy across 445 Terminal-Bench 2.1 trials; GPT-5.5 xhigh lifted from 83.1% to 92.1%; DeepSeek-V4 Flash from 82.7% to 88.1%; BusinessBench gains up to 10.04 points on some task families; and a cost comparison of roughly $15 against $574.68 for the GPT reference run. The authors say the approach transfers between models without touching weights.
What it means These are the authors’ own numbers on their own system and not yet independently replicated — but the claim is the interesting one either way: if most of a long-horizon failure rate is harness rather than model, the cheapest available win is in your own runtime.
Where it came from arXiv