🤖AI · 20282026-09-14
At Least One Leading AI Coding Agent Will Exceed An 85% Resolution Rate On The SWE-bench Verified Benchmark.2028

At Least One Leading AI Coding Agent Will Exceed An 85% Resolution Rate On The SWE-bench Verified Benchmark.

Will AI Coding Agents Surpass 85% On SWE-bench Verified By 2028?

A prediction states that at least one leading AI coding agent will achieve a solution rate above 85% on the SWE-bench Verified benchmark. The probability is estimated at 62%, with a target year of 2028. This article analyzes the trend, the counterarguments, and the verification criteria.

What Is The Current State Of SWE-bench Verified Scores?

As of February 2026, the public leaderboard at swebench.com shows Claude 4.5 Opus leading with a 76.8% solve rate. Gemini 3 Flash and MiniMax M2.5 follow closely at 75.8%, while DeepSeek V3.2 reaches 70.0%. This progress is not isolated to one lab. Multiple independent model families show simultaneous improvement, which strengthens the reliability of the trend.

How Fast Is The Improvement Rate On This Benchmark?

When OpenAI introduced SWE-bench Verified in mid-2024, top models scored around 60%. By February 2026, the leader reached 76.8%. This represents an annual gain of roughly 8 to 10 percentage points. If this pace continues, the 85% threshold becomes reachable by late 2028. Additionally, the cost per task for agents producing similar results has dropped below $0.10 for lower-tier models, accelerating enterprise adoption.

What Are The Main Counterarguments To This Prediction?

Benchmarks tend to saturate. Moving from 76% to 85% is harder than moving from 50% to 60% because the remaining problems are long-horizon, multi-step, and context-heavy. Benchmark contamination from training data leakage could also inflate scores. However, the concurrent progress across multiple independent model families, including Claude, Gemini, MiniMax, and DeepSeek, makes the 85% threshold plausible by 2028.

Why Does This Matter For Enterprise Software Development?

If the 85% threshold is crossed, most routine coding tasks in enterprise settings could be delegated to autonomous agents. This would shift developer roles toward oversight, architecture, and complex problem-solving. The cost reduction further accelerates this shift, making AI coding agents not only more capable but also economically viable at scale.

Frequently Asked Questions

How Is The 85% Threshold Verified?

The verification criterion is clear: at least one model must show a solve rate of 85% or higher on the public SWE-bench Verified leaderboard. This can be checked directly at swebench.com. No private or self-reported scores count.

Which Models Are Currently Leading The Benchmark?

As of February 2026, Claude 4.5 Opus leads with 76.8%, followed by Gemini 3 Flash at 75.8% and MiniMax M2.5 at 75.8%. DeepSeek V3.2 is at 70.0%. These figures are sourced from the public swebench.com leaderboard.

What Happens If The Prediction Fails?

If saturation slows progress, the 85% mark may slip beyond 2028. However, the current trajectory of multiple independent labs, combined with falling costs, suggests that even a delay would be short. The benchmark itself may also evolve, but the verified threshold remains the reference point for this prediction.

Probability
%62
Verification Criteria
At least one model shows an 85% or higher resolution rate on the public SWE-bench Verified leaderboard (verifiable via swebench.com).
Confidence Level
MediumResolution rose from roughly 60% in mid-2024 to 76.8% (Claude 4.5 Opus) by February 2026; an improvement of about +8-10 points per year makes 85% plausible by 2028, though saturation could slow it.
Loading…

Related Predictions

All predictions