
At Least One Leading AI Coding Agent Will Exceed An 85% Resolution Rate On The SWE-bench Verified Benchmark.
49fc88da3f9a0188 · Resolution source: swebench.com · SWE-bench official leaderboardAt Least One Leading AI Coding Agent Will Exceed An 85% Resolution Rate On The SWE-bench Verified Benchmark.
At Least One Leading AI Coding Agent Will Exceed An 85% Resolution Rate On The SWE-bench Verified Benchmark. Probability: 82%. Confidence Level: Medium.
Will AI Coding Agents Surpass 85% On SWE-bench Verified By 2028?
A prediction states that at least one leading AI coding agent will achieve a solution rate above 85% on the SWE-bench Verified benchmark. The probability is estimated at 62%, with a target year of 2028. This article analyzes the trend, the counterarguments, and the verification criteria.
What Is The Current State Of SWE-bench Verified Scores?
As of February 2026, the public leaderboard at swebench.com shows Claude 4.5 Opus leading with a 76.8% solve rate. Gemini 3 Flash and MiniMax M2.5 follow closely at 75.8%, while DeepSeek V3.2 reaches 70.0%. This progress is not isolated to one lab. Multiple independent model families show simultaneous improvement, which strengthens the reliability of the trend.
How Fast Is The Improvement Rate On This Benchmark?
When OpenAI introduced SWE-bench Verified in mid-2024, top models scored around 60%. By February 2026, the leader reached 76.8%. This represents an annual gain of roughly 8 to 10 percentage points. If this pace continues, the 85% threshold becomes reachable by late 2028. Additionally, the cost per task for agents producing similar results has dropped below $0.10 for lower-tier models, accelerating enterprise adoption.
What Are The Main Counterarguments To This Prediction?
Benchmarks tend to saturate. Moving from 76% to 85% is harder than moving from 50% to 60% because the remaining problems are long-horizon, multi-step, and context-heavy. Benchmark contamination from training data leakage could also inflate scores. However, the concurrent progress across multiple independent model families, including Claude, Gemini, MiniMax, and DeepSeek, makes the 85% threshold plausible by 2028.
Why Does This Matter For Enterprise Software Development?
If the 85% threshold is crossed, most routine coding tasks in enterprise settings could be delegated to autonomous agents. This would shift developer roles toward oversight, architecture, and complex problem-solving. The cost reduction further accelerates this shift, making AI coding agents not only more capable but also economically viable at scale.
Frequently Asked Questions
How Is The 85% Threshold Verified?
The verification criterion is clear: at least one model must show a solve rate of 85% or higher on the public SWE-bench Verified leaderboard. This can be checked directly at swebench.com. No private or self-reported scores count.
Which Models Are Currently Leading The Benchmark?
As of February 2026, Claude 4.5 Opus leads with 76.8%, followed by Gemini 3 Flash at 75.8% and MiniMax M2.5 at 75.8%. DeepSeek V3.2 is at 70.0%. These figures are sourced from the public swebench.com leaderboard.
What Happens If The Prediction Fails?
If saturation slows progress, the 85% mark may slip beyond 2028. However, the current trajectory of multiple independent labs, combined with falling costs, suggests that even a delay would be short. The benchmark itself may also evolve, but the verified threshold remains the reference point for this prediction.
Related Predictions
- Common Global Standards For AI Ethics And Governance Are Established. 2028 · AI
- AGI-like Systems Reach Human Expert Level In Narrow Domains Such As Medicine And Law. 2028 · AI
- AI-assisted Code Generation Shortens Software Development Time By Up To 40%. 2027 · AI
- Multimodal AI Models Process Text, Images, Audio And Video Simultaneously And Enter Mainstream Applications. 2027 · AI
- AGI Precursors Begin To Appear In Pilot Projects As Systems Approaching Human-level Performance In Narrow Domains. 2027 · AI
