Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

| Source: MarkTechPost

Tags: SWE-2, Cognition, Devin, Kimi K3, FrontierCode, coding models, reinforcement learning

Cognition's SWE-2 coding model, post-trained from Kimi K3 via RL, scores 50% on FrontierCode 1.1 — within 1 point of Fable 5.1 — while cutting costs 64% and running 81% cheaper with 58% fewer steps than predecessor SWE-1.7. Available exclusively inside Devin; no open weights or standalone API.

Details

Cognition has released SWE-2, its most capable coding model, built by post-training Kimi K3 — Moonshot AI's 2.8 trillion-parameter model — with reinforcement learning. The headline claim is 50.0% on FrontierCode 1.1 Main, Cognition's proprietary benchmark, placing it within one point of Fable 5.1 at 64% lower cost. The efficiency story is the most concrete data: SWE-2 medium outscores SWE-1.7 while taking 58% fewer turns and costing 81% less. Mean steps drop from 127 (SWE-1.7) to 53 at medium effort. The model leads Terminal-Bench 2.1 at 92.8% and scores 73.0% on DeepSWE 1.1. The weak spot is Terminal-Bench 4 at 27.3%, where Fable 5.1 (55.8%) and GPT-6 Astra (57.9%) hold a 30-point advantage. SWE-2 introduces three selectable reasoning-effort levels trained in a single RL run using Pareto-informed cost penalties. Behavioral improvements include stronger end-to-end test coverage, better resourcefulness when tools are blocked, and re-derivation under challenge rather than simple re-assertion. Key caveat: all benchmark comparisons against rivals use Cognition's own evaluation harness — independent verification is absent. SWE-2 has no open weights or standalone API; it runs exclusively inside Devin Desktop and CLI today, with Devin Web and Fusion rolling out soon.