Los modelos Claude de Anthropic lideran los benchmarks de programación SWE-bench

29 fuentes
  • A fecha del 8 de septiembre, los modelos Claude de Anthropic ocupan el primer puesto en los dos principales benchmarks de programación SWE-bench, superando a sus rivales de OpenAI y Google.
  • Las puntuaciones han pasado del 49% a principios de 2025 al 96% en SWE-bench Verified, que evalúa la capacidad de la IA para resolver problemas reales de GitHub de forma autónoma.
  • Estos resultados llegan en pleno auge de la inversión en programación mediante IA: Cognition acaba de recaudar 2.000 millones de dólares con una valoración de 48.000 millones para su asistente de programación Devin, según TechCrunch.
Fuentes (29)
  1. 1 SWE-bench Pro Leaderboard (September 2026): Claude Fable 5.1 ... benchlm.ai
  2. 2 SWE-bench Verified Leaderboard (September 2026): Top Scores benchlm.ai
  3. 3 Claude SWE-Bench Performance - Anthropic www.anthropic.com
  4. 4 Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market | TechCrunch techcrunch.com
  5. 5 SWE-bench Pro Leaderboard: Claude Fable 5.1 Takes #1 at 81.2% ... codingfleet.com
  6. 6 SWE-bench Scores and Leaderboard Explained (2026) dev.to
  7. 7 SWE-bench Verified Leaderboard 2026: Latest Coding Agent Scores leaderboard.steel.dev
  8. 8 Claude Opus 4.7 Benchmarks Explained - Vellum www.vellum.ai
  9. 9 From 80% to 93.9%: Why the Claude Mythos SWE-Bench ... www.mindstudio.ai
  10. 10 Claude Opus 4.8 on SWE-Bench Pro via AI Gateway www.truefoundry.com
  11. 11 Claude Sonnet 4.5 Tops SWE-Bench Verified, Extends Coding ... www.infoq.com
  12. 12 AWS Gives Students Free Year of Kiro AI Coding Tool www.techbuzz.ai
  13. 13 Claude Fable 5.1: Features, Benchmarks, and Pricing - DataCamp www.datacamp.com
  14. 14 Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic www.anthropic.com
  15. 15 The SWE-Bench Illusion: When State-of-the-Art LLMs Remember ... arxiv.org
  16. 16 Claude Fable 5.1 Benchmarks: Scores and What They Mean emergent.sh
  17. 17 IBM Chief: AI Should Enhance, Not Replace, Workforce www.chosun.com
  18. 18 SWE-bench February 2026 leaderboard update simonwillison.net
  19. 19 SWE-Bench Verified Leaderboard - LLM Stats llm-stats.com
  20. 20 SWE-bench Lite Leaderboard 2026 - Compare AI Model Scores pricepertoken.com
  21. 21 LLM Coding Leaderboard - SWE-bench / LiveCodeBench / SWE ... www.datalearner.com
  22. 22 SWE-rebench Leaderboard (March, April and May 2026) - Reddit www.reddit.com
  23. 23 Claude Benchmarks (2026): Opus 5, Sonnet 5, and Fable 5 at 95 ... www.morphllm.com
  24. 24 SWE-bench Verified www.swebench.com
  25. 25 Anthropic: Claude Now Writes Over 80% of Its Production Code aitoolsrecap.com
  26. 26 Live blog: Code w/ Claude 2026 simonwillison.net
  27. 27 Anthropic Release Notes - September 2026 Latest Updates releasebot.io
  28. 28 Anthropic 2026: Every Claude Model, Agent & Tool linas.substack.com
  29. 29 Introducing Claude Opus 4.5 - Anthropic www.anthropic.com