Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

benchlm+1anthropic+1techcrunchAnthropic's Claude AI models hold the top positions across every major variant of the SWE-bench coding benchmark as of September 2026, with Claude Opus 5 reaching 96% on SWE-bench Verified and Claude Fable 5.1 scoring 81.2% on the more demanding SWE-bench Pro. The results underscore Anthropic's widening lead in AI-assisted software engineering, even as rivals from OpenAI and Google Alphabet Inc. continue to close gaps on other fronts.benchlm+2
SWE-bench, created by Princeton University researchers, tests whether AI models can resolve real-world GitHub issues by generating patches that pass both new and existing test suites. On SWE-bench Verified, a curated subset of 500 problems, Claude Opus 5 leads at 96%, followed by Claude Mythos 5 at 95.5% and Claude Fable 5 at 95%. On SWE-bench Pro, which spans 1,865 problems across 41 repositories and tests longer-horizon engineering tasks, Claude Fable 5.1 tops the leaderboard at 81.2%.benchlm+4
The trajectory has been steep. In early 2025, Anthropic's Claude 3.5 Sonnet scored 49% on SWE-bench Verified. By late 2025, Claude Sonnet 4.5 reached 77.2%. Claude Opus 4.5 pushed that to 80.9% by early 2026, and each subsequent model generation — Opus 4.7, Opus 4.8, Mythos, and the Opus 5 and Fable 5 families — extended the lead further.dev+5
Anthropic's dominance on SWE-bench comes amid an intensifying AI coding arms race. Cognition, the startup behind the Devin coding assistant, just raised $2 billion at a $48 billion valuation, with annualized revenue growing from $492 million to $900 million in four months. Amazon Web Services Amazon.com, Inc. is pushing its Kiro coding tool to students at 132 universities worldwide. OpenAI's GPT-5.6 Sol, its latest flagship, has drawn roughly level with Claude on some general benchmarks but trails on agentic coding tasks, scoring 37.3% on Terminal-Bench 4.0 compared to Fable 5.1's 55.8%.techcrunch+3
Not everyone treats SWE-bench scores as definitive. Research published on arXiv found that top models could identify buggy file paths with up to 76% accuracy using only issue descriptions, raising questions about whether models are genuinely reasoning or partially relying on patterns in the data. Anthropic itself has shifted its headline benchmarks for newer models toward longer-horizon agentic tasks like Terminal-Bench, where Fable 5.1 more than doubled its predecessor's score on scientific research problems. The company did not headline a new SWE-bench number for Fable 5.1 at all, suggesting even Anthropic views the benchmark as approaching saturation at the top end.arxiv+2