2028At Least One Leading AI Coding Agent Will Exceed An 85% Resolution Rate On The SWE-bench Verified Benchmark.
Will AI Coding Agents Surpass 85% On SWE-bench Verified By 2028?
A prediction states that at least one leading AI coding agent will achieve a solution rate above 85% on the SWE-bench Verified benchmark. The probability is estimated at 62%, with a target year of 2028. This article analyzes the trend, the counterarguments, and the verification criteria.
What Is The Current State Of SWE-bench Verified Scores?
As of February 2026, the public leaderboard at swebench.com shows Claude 4.5 Opus leading with a 76.8% solve rate. Gemini 3 Flash and MiniMax M2.5 follow closely at 75.8%, while DeepSeek V3.2 reaches 70.0%. This progress is not isolated to one lab. Multiple independent model families show simultaneous improvement, which strengthens the reliability of the trend.
How Fast Is The Improvement Rate On This Benchmark?
When OpenAI introduced SWE-bench Verified in mid-2024, top models scored around 60%. By February 2026, the leader reached 76.8%. This represents an annual gain of roughly 8 to 10 percentage points. If this pace continues, the 85% threshold becomes reachable by late 2028. Additionally, the cost per task for agents producing similar results has dropped below $0.10 for lower-tier models, accelerating enterprise adoption.
What Are The Main Counterarguments To This Prediction?
Benchmarks tend to saturate. Moving from 76% to 85% is harder than moving from 50% to 60% because the remaining problems are long-horizon, multi-step, and context-heavy. Benchmark contamination from training data leakage could also inflate scores. However, the concurrent progress across multiple independent model families, including Claude, Gemini, MiniMax, and DeepSeek, makes the 85% threshold plausible by 2028.
Why Does This Matter For Enterprise Software Development?
If the 85% threshold is crossed, most routine coding tasks in enterprise settings could be delegated to autonomous agents. This would shift developer roles toward oversight, architecture, and complex problem-solving. The cost reduction further accelerates this shift, making AI coding agents not only more capable but also economically viable at scale.
Frequently Asked Questions
How Is The 85% Threshold Verified?
The verification criterion is clear: at least one model must show a solve rate of 85% or higher on the public SWE-bench Verified leaderboard. This can be checked directly at swebench.com. No private or self-reported scores count.
Which Models Are Currently Leading The Benchmark?
As of February 2026, Claude 4.5 Opus leads with 76.8%, followed by Gemini 3 Flash at 75.8% and MiniMax M2.5 at 75.8%. DeepSeek V3.2 is at 70.0%. These figures are sourced from the public swebench.com leaderboard.
What Happens If The Prediction Fails?
If saturation slows progress, the 85% mark may slip beyond 2028. However, the current trajectory of multiple independent labs, combined with falling costs, suggests that even a delay would be short. The benchmark itself may also evolve, but the verified threshold remains the reference point for this prediction.
Related Predictions
- By The End Of 2027, Global Electricity Consumption Of AI Servers Will Exceed The Total Consumption Of Conventional (Non-AI) Data Center Hardware. 2027 · AI
- US Data Center Electricity Demand Will Exceed The 60 Gigawatt Threshold By The End Of 2027 (Up From ~31 GW In 2025). 2027 · AI
- The EU's Compliance Deadline For High-risk AI Systems, Deferred From August 2026 To December 2027, Will Come Into Force On 2 December 2027 Without Being Postponed Again. 2027 · AI
- By The End Of 2027, An Independent Enterprise Research Report Will State That At Least 15% Of Large Companies Have Delegated At Least One End-to-end Business Process To An AI Agent, Mostly Without Human Intervention. 2027 · AI
- By The End Of 2027, A Leading AI Lab Will Announce That It Has Reached At Least 50% Accuracy On The Hardest Tier (Research-level) Of The FrontierMath Benchmark. 2027 · AI
