Cognition SWE-2 in Devin: Near-Fable Scores at 64% Lower Cost

Coffee Summary

  • FACT: Cognition launched SWE-2 on 10 Sep 2026 in Devin Desktop and CLI; Web and Fusion rollout is underway. No standalone API or open weights.
  • CLAIM: Cognition reports 50.0% on FrontierCode 1.1 Main versus Fable 5.1 at 50.9%, with 64% lower cost; DeepSWE 1.1 is reported at 73.0%.
  • FACT: Terminal-Bench 4 is a counter-signal: SWE-2 27.3% versus Fable 55.8% and GPT-6 Astra 57.9%.

What happened

Cognition positions SWE-2 as a Pareto model for Devin, post-trained from Moonshot Kimi K3. It jointly trains medium, high, and max effort with a cost-penalized reward. SWE-2 medium reportedly makes its first real edit after 18 steps versus 48 for SWE-1.7.

Why it matters

For Devin customers, the question is frontier-level solve rate versus cost per successful task. SWE-2 may be a viable default for scoped tickets, but the Terminal-Bench gap means buyers should not assume parity on hard agentic shell work.

Buyer checklist

  1. Run a 20–50 ticket pilot at medium effort and log turns, failures, dollars, and ACUs.
  2. Include multi-step terminal and infrastructure-debug tasks.
  3. Keep a frontier API escape hatch until your own harness confirms parity.
  4. Do not plan a custom CI integration until Cognition ships a standalone API.

AIImpish Take

SWE-2 is primarily a Devin product story: near-Fable scores at a claimed steep discount, with a clear Terminal-Bench 4 hole. Pilot it inside Devin if you already use the platform; choose portable frontier APIs when you need an endpoint.