Dashboard

Cognition SWE-2: Near-Frontier Coding, 64% Cheaper

SWE-2 lands within a point of Fable 5.1 on Cognition's mergeability benchmark at a fraction of the cost. The benchmark, the base model, and the caveats.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
11 September 20261 min read

Cognition SWE-2: Near-Frontier Coding, 64% Cheaper

Cognition released SWE-2 on 10 September 2026, and the headline is a trade rather than a win. On Cognition's FrontierCode 1.1 Main benchmark, SWE-2 scores 50.0 percent against Fable 5.1's 50.9 percent, and Cognition says it costs 64 percent less to run. Nine tenths of a benchmark point for roughly a third of the bill is a trade most people building on their own money should at least look at.

The scoreboard

FrontierCode 1.1 Main, as published by Cognition:

  • GPT-6 Astra: 53.3 percent

  • Fable 5.1: 50.9 percent

  • SWE-2: 50.0 percent

  • Grok 4.6: 48.0 percent

  • GPT-5.6 Sol: 47.5 percent

  • SWE-1.7: 42.0 percent

The benchmark itself is worth a sentence, because it is not another pass-the-unit-test score. FrontierCode asks whether a human maintainer would actually merge the pull request the model wrote. That is a closer proxy for the thing you care about than a green test suite, since an agent that edits tests until they pass scores perfectly on the wrong question.

What it is built on

SWE-2 is post-trained from Kimi K3, Moonshot's 2.8-trillion-parameter model, which Cognition notes had already been through extensive reinforcement learning for agentic coding before they started. So this is a lab taking someone else's coding-tuned base and specialising it further, not a from-scratch frontier run. Worth knowing if you want to understand what post-training actually does to a model.

Cognition also ships three effort levels, medium, high and max, and says it trained all three in a single run rather than producing separate models per setting. On FrontierCode 1.1 Main, the medium setting beats the older SWE-1.7 while taking 58 percent fewer turns and costing 81 percent less on average.

Where the cost claim lands

Two figures were published: 64 percent cheaper than Fable 5.1, and roughly a quarter the price of GPT-6 Astra. Fewer turns is where a lot of that comes from, and it compounds in a way per-token pricing tables hide. An agent that reaches the same patch in 40 percent of the turns is not just cheaper, it is faster to review and less likely to wander into files nobody asked it to touch.

The caveat is the usual one and it is not small. These are the vendor's numbers, on the vendor's benchmark, about the vendor's model. FrontierCode is a genuinely interesting benchmark and Cognition built it, which are both true at once. If SWE-2's price is what draws you, the number that decides it is what an agent costs you per month on your own work, not a published multiple.

The open-base detail nobody is leading with

Kimi K3 is an open-weights model. A second company took it, spent its own reinforcement learning budget on it, and shipped a product that competes with closed frontier models on a coding benchmark. That is a different supply chain from the one most people are planning around, and it has a practical consequence: the ceiling on what a small lab can ship without training a frontier model itself went up again.

For buyers it cuts two ways. More vendors at the frontier means more price pressure, which is good. It also means the model behind your coding agent may now be three companies deep, which matters the day one of them changes terms or deprecates a tier. If that sounds abstract, it is the same dependency shape that made what to do when an AI model gets deprecated a question worth having an answer to.

Should you switch

Probably not on a launch post. The pattern worth keeping is boring: run your own small evaluation set before moving anything real, because a model that ranks a point lower on a public leaderboard can rank first on your codebase, and the reverse happens just as often. Our walkthrough of evaluating a new model release before switching has the rubric.

SWE-2 is available now in Devin Desktop and the Devin CLI, with Devin Web and Fusion to follow. If you are already paying Cognition, the medium setting is the cheap experiment. If you are not, the more interesting signal is structural: a 2.8-trillion-parameter open base, post-trained by a second company, landing within a point of a frontier lab's coding model. That gap closing is the story, and it will not be the last time it happens this year.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.