Retrieved article excerpt
Open article · Retrieved 2026-09-10T16:24:25.303205+00:00
Introducing SWE-2: Pushing the Pareto Frontier By The Cognition Team 09.10.26 Today we’re introducing SWE-2, our most advanced coding model yet. It pushes the Pareto frontier of capability and cost, achieving 50.0% on FrontierCode 1.1 Main 1 , within one point of Fable 5.1 while being 64% cheaper. With SWE-2, we scaled RL to the multi-trillion-parameter regime for the first time, building on the SWE-1.7 2 training infrastructure and recipe. The key addition is an RL algorithm that trains all reasoning-effort levels in a single run, advancing the whole cost–performance frontier. base model end of training The result is our closest model yet to the frontier. On FrontierCode 1.1 Main and DeepSWE 1.1, SWE-2 beats SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price, and comes within a few points of GPT-6 Astra at a quarter of the cost. See how models rank on the FrontierCode leaderboard → SWE-2 is post-trained from Kimi K3 3 , a 2.8T-parameter model that had already undergone extensive RL for agentic coding. As with SWE-1.7, our RL still finds substantial headroom, adding 5–6 points on many benchmarks and shifting K3’s entire cost–performance frontier. Coding benchmark results Benchmark SWE-2 Kimi K3 Grok 4.6 Fable 5.1 GPT-5.6 Sol GPT-6 Astra SWE-1.7 FrontierCode 1.1 Main 50.0 % 44.2 % 48.0 % 50.9 % 47.5 % 53.3 % 42.0 % DeepSWE 1.1 73.0 % 68.5 % 67.5 % 67.4 % 72.7 % 74.1 % 37.7 % Terminal-Bench 2.1 92.8 % 88.3 % 88.4 % 91.4 % 88.8 % 89.9 % 81.5 % Terminal-Bench 4 27.3 % 21.5 % 20.3 % 55.8 % 37.3 % 57.9 % 7.6 % The rest of this post covers what SWE-2 does differently and how we trained it. We begin with SWE-2’s behavior, focusing on the characteristics that make it more efficient and intelligent compared to our previous models. Then, we detail the post-training advances behind SWE-2: Cost penalties. We apply a linear cost penalty per effort level in a single RL run, with each penalty tuned to the local slope of the base model’s Pareto frontier. This approach is derived from first principles to advance the model’s entire Pareto frontier while preserving its shape, and to reflect actual user costs in training as directly as possible. Reward baselines. We derive the length-weighted reward baseline we have used since SWE-1.6 and show how it significantly stabilizes training. RL rollout serving. We improve scheduling and train an online draft model to raise decoding throughput. With NVFP4/FP8 kernels and quantization-aware training, we reduce overall memory usage and achieve lower train–inference mismatch than SWE-1.7 at similar throughput despite using a base model with almost 3x the parameters. Training data. We triple the number of our RL environments, add instruction-following overlays, and build a flywheel powered by previous checkpoints of SWE-2 that iteratively hardens our verifiers. SWE-2 is available starting today in Devin Desktop and CLI . We’re also rolling it out on Devin Web and Fusion . Model Behavior # SWE-2’s improvements in intelligence and efficiency are closely connected. Stronger engineering judgment allows the agent to write more complete solutions alongside fewer detours and redundant reads. On FrontierCode 1.1 Main, we see that SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average. SWE-1.7 vs. SWE-2 on FrontierCode 1.1 Main: Mean Steps per run SWE-1.7 127 SWE-2 medium 53 SWE-2 high 80 SWE-2 max 98 0 50 100 Explore (read / grep / ls) Plan / todo Write / edit code Build (make / lint) Run tests git add / commit Final message Mean metric over all 100-task FrontierCode 1.1 Main tasks, using three runs per task per model and grouped by the tools each step calls. In our previous post 2 , we observed SWE-1.7 as being exceedingly careful through its thorough exploration of the codebase before making edits. While boosting performance, this led to user feedback that SWE-1.7 tended to over-explore and overthink on simple tasks. Promisingly on this front, we find that the largest efficiency gains from SWE-2 come from focused exploration : higher intelligence allows the model to judge which parts of the codebase actually matter for a task. This allows SWE-2 to begin implementation sooner: on FrontierCode 1.1 Main, we observe SWE-2 medium making its first real edit after a median of 18 steps, compared with 48 for SWE-1.7. Example trajectories Show From testing SWE-2 internally, we observed that the higher model capabilities also manifested in the following behavioral patterns: Test coverage: SWE-2 is better at writing tests that check an implementation end-to-end, catching regressions and edge cases more reliably. Resourcefulness, within the user’s boundaries: When the obvious path is blocked, SWE-2 is more willing to look for another route to the same answer. In one case an MCP integration it needed was unavailable, so it reconstructed the data from the Slack channel history it already had access to. Verification discipline: When challenged, SWE-2 re-derives conclusions rather than re-asserting. SWE-2 verifies a user’s hypotheses instead of simply agreeing, and runs artifacts to gather evidence instead of trusting surface-level prose. The result is a model whose conclusions you can trust. We observe real behavioral differences between effort levels as well. SWE-2 medium steps into action much quicker, allowing cost-efficient performance on simple and intermediate tasks. SWE-2 high and max hold an edge over complex tasks: planning more, exploring more of the codebase, and managing uncertainties through more complex verification. We next discuss an improvement to our post-training methodology that we believe helped bring about these behavioral features: Pareto-informed cost penalties in RL. Pushing the Pareto Frontier with RL # As models become more intelligent and expensive, cost–performance tradeoffs grow increasingly important in the coding agent landscape. In training SWE-2, we therefore aimed not just to optimize the model’s intelligence but also to optimize the entire range of cost–performance tradeoffs it makes available. Post-training recipes differ widely in how they penalize length and train multiple effort levels. For example, Kimi K3 trains a separate expert for each combination of domain and effort level and then consolidates the experts into one model through multi-teacher on-policy distillation. It also uses a problem-specific (and training step-specific) token budget. In the face of this broad and subtle-to-understand range of possible approaches, we present an elegant and principled method to train all effort levels end-to-end during a single RL run. Progress of the Pareto frontier during training Kimi K3 end of training We accomplish this by using a cost-penalized reward function of the form R = S − λ e C , R=S-\lambda_e C, R = S − λ e C , where S ∈ { 0 , 1 } S \in \{0,1\} S ∈ { 0 , 1 } denotes whether a rollout was successful, C C C denotes the cost of a rollout (a mix of inference cost in USD and rollout time), e e e denotes the effort level, and λ e \lambda_e λ e is a parameter tuned to match the slope of the Pareto curve of the base model at effort level e e e . Approximating the Pareto curve tangents of Kimi K3 These choices might seem counterintuitive, but as we will now see, they are logical conclusions derived from our goal of pushing the Pareto frontier. Deriving the Cost Penalty # We next explain how we chose an RL objective R R R that directly optimizes the model’s cost–performance Pareto frontier. Here, “cost” refers to average cost and “performance” refers to solve rate, both averaged over a distribution D \mathcal D D of training tasks. Recall that points on the cost–performance plane depend on the task distribution’s average cost and average solve rate but otherwise do not depend on D \mathcal D D . Therefore, to align the RL objective with a model’s position in the plane, we want the expectation of R R R over D \mathcal D D to depend only on this average cost and solve rate. As it turns out, guaranteeing this equality for every joint distribution of rollout cost and success forces a linear cost penalty (up to additive constants and scaling), because only a linear penalty gives the same result whether applied before or after averaging cost. For the interested reader, we prove this claim rigorously in Appendix B . Now that we have our reward function R = S − λ e C R=S-\lambda_e C R = S − λ e C , the final task is selecting λ e \lambda_e λ e for each effort level. While setting λ e \lambda_e λ e might at first feel like a hyperparameter optimization problem, it turns out that our goal of pushing the Pareto frontier upwards again dictates how we should make this choice. Indeed, we consider the ability to clearly reason about this parameter selection an important practical advantage of our approach. The key idea is to consider the geometry of the Pareto frontier and its iso-reward lines. To do so, fix an effort level and let ( c , s ) (c,s) ( c , s ) be the corresponding point on the current frontier, with average reward J = s − λ e c J=s-\lambda_e c J = s − λ e c . Its iso-reward line satisfies s = λ e c + J s=\lambda_e c+J s = λ e c + J , and therefore has slope λ e \lambda_e λ e . In the left panel below, we see a failure case where λ high \lambda_\text{high} λ high is set too large: the model is rewarded for performing an unhelpful update, one where the model at high-effort starts to behave like the medium-effort version. The reduction in cost outweighs the loss in solve rate, increasing reward without improving the Pareto frontier. In the right panel, λ high \lambda_\text{high} λ high matches the frontier’s slope at the current high-effort point. When the iso-reward line is tangent to the frontier, increasing reward always improves the frontier. We can formalize this geometrical intuition with a bit of algebra. Let m m m be the local slope of the Pareto frontier at ( c , s ) (c,s) ( c , s ) . A small movement along the frontier changes the solve rate by Δ s ≈ m Δ c \Delta s\approx m\Delta c Δ s ≈ m Δ c , so the corresponding change in average reward is Δ J = Δ s − λ e Δ c ≈ ( m − λ e ) Δ c . \Delta J = \Delta s - \lambda_e \Delta c \approx (m - \lambda_e)\Delta c. Δ J = Δ s − λ e Δ c ≈ ( m − λ e ) Δ c . Thus, letting λ e = m \lambda_e = m λ e = m ensures that the objective J J J is unaffected (to first order) by movements along the Pareto curve. Length-Weighted Reward Baseline # We’re also sharing the reward baseline we’ve used since SWE-1.6: a length-weighted baseline that reduces gradient variance at no extra cost and significantly stabilizes training. Given a fixed prompt x x x and a group of n n n rollouts y 1 , … , y n y_1,\ldots,y_n y 1 , … , y n , the on-policy gradient estimator with baseline b b b is g ^ = 1 n ∑ i = 1 n ( R i − b ) ∇ θ log π θ ( y i ∣ x ) . \widehat g = \frac{1}{n}\sum_{i=1}^{n}(R_i-b)\,\nabla_\theta\log\pi_\theta(y_i\mid x). g = n 1 i = 1 ∑ n ( R i − b ) ∇ θ lo g π θ ( y i ∣ x ) . A reasonable proxy for reducing the gradient estimator’s variance is to minimize E [ ( R i − b ) 2 ] \mathbb E[(R_i-b)^2] E [( R i − b ) 2 ] . This gives the mean-reward baseline b = E [ R i ] b = \mathbb E[R_i] b = E [ R i ] , which in practice we estimate using the group baseline 4 b = 1 n ∑ i = 1 n R i b = \frac{1}{n}\sum_{i=1}^{n}R_i b = n 1 ∑ i = 1 n R i . Its dependence on the sampled rollouts introduces some bias in the gradient estimator, but this bias decays as 1 / n 1/n 1/ n and is small for large groups. We instead attempt to minimize the variance of the full gradient estimator g ^ \hat g g ^ . Following Greensmith, Bartlett, and Baxter (2004) 5 , 6 , the optimal baseline is b ⋆ = E [ R i ∥ ∇ θ log π θ ( y i ∣ x ) ∥ 2 ] E [ ∥ ∇ θ log π θ ( y i ∣ x ) ∥ 2 ] . b^\star = \frac{\mathbb E\left[R_i\left\|\nabla_\theta\log\pi_\theta(y_i\mid x)\right\|^2\right]}{\mathbb