Cognition released SWE-2 on September 10, calling it its most capable coding model to date. It scores 50.0 percent on the company's FrontierCode 1.1 Main benchmark, putting it within one point of Fable 5.1, a model generally placed at the frontier. Reaching that level costs 64 percent less, according to Cognition. The announcement is built less around raw benchmark numbers than around how far the balance between capability and cost can be pushed.
Shifting the target from higher scores to cheaper scores
The framing that runs through the announcement is how far the boundary connecting cost and capability has been moved outward. Cut spending and the score drops; raise the score and spending climbs. Cognition presents its result as moving that boundary itself, rather than lining up absolute scores in isolation.
The specific comparisons go like this. On both FrontierCode 1.1 Main and DeepSWE 1.1, SWE-2 beats the previous SWE-1.7 and Grok 4.6 on score and cost at the same time. It matches GPT-5.6 Sol, Fable 5 and Fable 5.1 on score while costing far less, and comes within a few points of the top-end GPT-6 Astra at a quarter of the price. For reference, Fable 5.1 at its medium setting is listed at 50.9 percent and 3.28 USD per task (about 500 yen).
1 USD = 154 JPY (as of September 10, 2026)
Fewer detours are where the savings come from
The clearest part of the cost story is a change in how the model works through a task. SWE-2 at its medium setting scores higher than SWE-1.7 while using 58 percent fewer turns and costing 81 percent less on average.
SWE-1.7 had a habit of reading widely around a codebase before touching any code, and users reported that it over-explored even on simple work. SWE-2 is better at judging which parts of a codebase actually matter, and it now reaches its first real edit after a median of 18 steps. SWE-1.7 needed 48, so the gap is more than half.
Cognition also reports changes in judgment quality: tests that exercise an implementation end to end, a willingness to find another route when the obvious path is blocked, and re-deriving conclusions from evidence when challenged rather than repeating the same claim. Speed here did not come at the price of care, is the argument.
The base is Kimi K3, a 2.8 trillion parameter model
What stands out is that SWE-2 was not built from scratch. It is post-trained from Kimi K3, the 2.8 trillion parameter model released by Moonshot AI and the largest open-weight model available. The starting point had already been through reinforcement learning for agentic coding, and Cognition trained further on top of it.
That additional work is credited with 5 to 6 points on many benchmarks, which suggests there is still room to move the cost-capability boundary even when starting from a public model. It also hints at a path for organizations that cannot afford to pre-train a giant model of their own but can invest in focused post-training.
Training every effort level in a single run
At the center of the cost reduction is a method that trains several reasoning-effort settings together in one reinforcement learning run.
The reward function carries a penalty proportional to the cost a rollout incurred, and the strength of that penalty varies by effort level. The strength is not tuned by trial and error; it is matched to the slope of the base model's cost-capability curve at that point. If the penalty is too strong, a high-effort setting starts behaving like a lower one, spending falls, and the reward rises without anything actually improving. Matching the slope keeps rising reward aligned with an outward move of the boundary.
The training substrate was widened as well. Cognition tripled the number of reinforcement learning environments, added training that keeps multiple instructions in view without losing the underlying task, and hardened its graders against reward hacking using earlier checkpoints of SWE-2 itself. On the serving side, an improved speculative decoding setup lengthened accepted sequences by 15 percent, and batching incoming prefill requests raised throughput per GPU by 10 to 20 percent.
Checks on building from a Chinese open model
Building on open weights raises questions about bias in the output, and Cognition published two evaluations on that point.
The first sends 145 politically sensitive questions about China in English, Simplified Chinese and Traditional Chinese, and checks whether the model adopts the official position as its own. SWE-2 passed 98.0 percent overall: 99.8 percent in English, 95.2 percent in Simplified Chinese and 99.1 percent in Traditional Chinese.
The second tests whether the identity of the requester or the language of the request makes a model more willing to implement dangerous functionality. Across six models, SWE-2, Kimi K3, GLM 5.3, GPT 5.6, Fable 5.1 and Opus 5, no framing produced a statistically significant difference.
Available through Devin first
SWE-2 is available from launch in Devin Desktop and the command line client, with Devin Web and Fusion following. Rather than being sold broadly as a standalone API, it replaces the engine inside Cognition's own coding agent products. Because cost moves sharply with the effort setting, practical use assumes matching the setting to the difficulty of the work.
How to read the numbers
There are a few things worth checking before taking the figures at face value. FrontierCode, the main axis of comparison, is a benchmark Cognition built, and models without published results were measured in Cognition's own evaluation harness. A model scoring well on a metric shaped around its maker's own workload is not surprising, so independent reproduction is the thing to watch next.
The cost comparison is also explicitly based on list pricing. Actual spending shifts with discounts, usage volume and which effort setting becomes the default. Measuring which setting carries most of your own work is the practical first step before comparing prices.
Summary
Cognition released SWE-2 on September 10 with a score of 50.0 percent on FrontierCode 1.1 Main, staying within one point of frontier-tier models while cutting cost by 64 percent. Against the previous generation, turns fell 58 percent and average cost fell 81 percent. It is post-trained from Moonshot AI's Kimi K3, using reinforcement learning that handles every effort level in one run to push the cost-capability boundary outward. The release points to a route toward frontier-level results without building a giant model in-house.
