Modely Claude od Anthropic ovládly žebříčky programování SWE-bench

29 zdroje
  • Modely Claude od společnosti Anthropic drží k 8. září první místo v obou hlavních programovacích testech SWE-bench a překonávají rivaly z OpenAI a Googlu.
  • Skóre vzrostlo ze 49 % na začátku roku 2025 na 96 % v testu SWE-bench Verified, který prověřuje schopnost AI samostatně řešit reálné problémy na GitHubu.
  • Výsledky přicházejí v době boomu investic do AI programování: společnost Cognition podle TechCrunch právě získala 2 miliardy dolarů při ocenění 48 miliard dolarů za svého programovacího asistenta Devin.
Zdroje (29)
  1. 1 SWE-bench Pro Leaderboard (September 2026): Claude Fable 5.1 ... benchlm.ai
  2. 2 SWE-bench Verified Leaderboard (September 2026): Top Scores benchlm.ai
  3. 3 Claude SWE-Bench Performance - Anthropic www.anthropic.com
  4. 4 Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market | TechCrunch techcrunch.com
  5. 5 SWE-bench Pro Leaderboard: Claude Fable 5.1 Takes #1 at 81.2% ... codingfleet.com
  6. 6 SWE-bench Scores and Leaderboard Explained (2026) dev.to
  7. 7 SWE-bench Verified Leaderboard 2026: Latest Coding Agent Scores leaderboard.steel.dev
  8. 8 Claude Opus 4.7 Benchmarks Explained - Vellum www.vellum.ai
  9. 9 From 80% to 93.9%: Why the Claude Mythos SWE-Bench ... www.mindstudio.ai
  10. 10 Claude Opus 4.8 on SWE-Bench Pro via AI Gateway www.truefoundry.com
  11. 11 Claude Sonnet 4.5 Tops SWE-Bench Verified, Extends Coding ... www.infoq.com
  12. 12 AWS Gives Students Free Year of Kiro AI Coding Tool www.techbuzz.ai
  13. 13 Claude Fable 5.1: Features, Benchmarks, and Pricing - DataCamp www.datacamp.com
  14. 14 Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic www.anthropic.com
  15. 15 The SWE-Bench Illusion: When State-of-the-Art LLMs Remember ... arxiv.org
  16. 16 Claude Fable 5.1 Benchmarks: Scores and What They Mean emergent.sh
  17. 17 IBM Chief: AI Should Enhance, Not Replace, Workforce www.chosun.com
  18. 18 SWE-bench February 2026 leaderboard update simonwillison.net
  19. 19 SWE-Bench Verified Leaderboard - LLM Stats llm-stats.com
  20. 20 SWE-bench Lite Leaderboard 2026 - Compare AI Model Scores pricepertoken.com
  21. 21 LLM Coding Leaderboard - SWE-bench / LiveCodeBench / SWE ... www.datalearner.com
  22. 22 SWE-rebench Leaderboard (March, April and May 2026) - Reddit www.reddit.com
  23. 23 Claude Benchmarks (2026): Opus 5, Sonnet 5, and Fable 5 at 95 ... www.morphllm.com
  24. 24 SWE-bench Verified www.swebench.com
  25. 25 Anthropic: Claude Now Writes Over 80% of Its Production Code aitoolsrecap.com
  26. 26 Live blog: Code w/ Claude 2026 simonwillison.net
  27. 27 Anthropic Release Notes - September 2026 Latest Updates releasebot.io
  28. 28 Anthropic 2026: Every Claude Model, Agent & Tool linas.substack.com
  29. 29 Introducing Claude Opus 4.5 - Anthropic www.anthropic.com