Beating AI News Flash: Cognition, the parent company of Coding Agent Devin, has released a new coding model, SWE-2. It directly builds on Moonshot AI's 2.8T-parameter Kimi K3 for continued reinforcement learning, and after further training, its scores on multiple coding benchmarks improved by about 5 to 6 percentage points.
On Cognition's own FrontierCode 1.1 Main, SWE-2 scored 50.0%, surpassing GPT-5.6 Sol's 47.5% and Grok 4.6's 48.0%. It trails Fable 5.1's 50.9% by only 0.9 percentage points, but costs 64% less. GPT-6 Astra scored 53.3%, with SWE-2 behind it by 3.3 percentage points, at about one-quarter of the cost.
SWE-2 also takes far fewer detours than the previous generation. At medium reasoning effort, it reduces the number of interaction turns needed to complete tasks by 58%, with average costs dropping 81%. However, on the harder Terminal-Bench 4, SWE-2 scores only 27.3%, significantly below Astra's 57.9% and Fable 5.1's 55.8%, so it is not yet fair to say it has fully caught up with frontier models.
SWE-2 is now available on Devin Desktop and CLI, and is beginning to roll out to Devin Web and Fusion.

