Claude-líkön Anthropic raða sér í efstu sæti SWE-bench forritunarprófa

29 heimildir
  • Claude-líkön Anthropic eru í efsta sæti á báðum helstu SWE-bench forritunarprófunum frá og með 8. september og skjóta keppinautum frá OpenAI og Google ref fyrir rass.
  • Einkunnir hafa rokið upp úr 49% snemma árs 2025 í 96% á SWE-bench Verified, sem prófar getu gervigreindar til að leysa raunveruleg GitHub-vandamál sjálfstætt.
  • Niðurstöðurnar koma á sama tíma og fjárfestingar í gervigreindri forritun aukast hratt: Cognition safnaði nýverið 2 milljörðum dala á 48 milljarða dala virði fyrir Devin-forritunaraðstoðarmann sinn, samkvæmt TechCrunch.
Heimildir (29)
  1. 1 SWE-bench Pro Leaderboard (September 2026): Claude Fable 5.1 ... benchlm.ai
  2. 2 SWE-bench Verified Leaderboard (September 2026): Top Scores benchlm.ai
  3. 3 Claude SWE-Bench Performance - Anthropic www.anthropic.com
  4. 4 Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market | TechCrunch techcrunch.com
  5. 5 SWE-bench Pro Leaderboard: Claude Fable 5.1 Takes #1 at 81.2% ... codingfleet.com
  6. 6 SWE-bench Scores and Leaderboard Explained (2026) dev.to
  7. 7 SWE-bench Verified Leaderboard 2026: Latest Coding Agent Scores leaderboard.steel.dev
  8. 8 Claude Opus 4.7 Benchmarks Explained - Vellum www.vellum.ai
  9. 9 From 80% to 93.9%: Why the Claude Mythos SWE-Bench ... www.mindstudio.ai
  10. 10 Claude Opus 4.8 on SWE-Bench Pro via AI Gateway www.truefoundry.com
  11. 11 Claude Sonnet 4.5 Tops SWE-Bench Verified, Extends Coding ... www.infoq.com
  12. 12 AWS Gives Students Free Year of Kiro AI Coding Tool www.techbuzz.ai
  13. 13 Claude Fable 5.1: Features, Benchmarks, and Pricing - DataCamp www.datacamp.com
  14. 14 Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic www.anthropic.com
  15. 15 The SWE-Bench Illusion: When State-of-the-Art LLMs Remember ... arxiv.org
  16. 16 Claude Fable 5.1 Benchmarks: Scores and What They Mean emergent.sh
  17. 17 IBM Chief: AI Should Enhance, Not Replace, Workforce www.chosun.com
  18. 18 SWE-bench February 2026 leaderboard update simonwillison.net
  19. 19 SWE-Bench Verified Leaderboard - LLM Stats llm-stats.com
  20. 20 SWE-bench Lite Leaderboard 2026 - Compare AI Model Scores pricepertoken.com
  21. 21 LLM Coding Leaderboard - SWE-bench / LiveCodeBench / SWE ... www.datalearner.com
  22. 22 SWE-rebench Leaderboard (March, April and May 2026) - Reddit www.reddit.com
  23. 23 Claude Benchmarks (2026): Opus 5, Sonnet 5, and Fable 5 at 95 ... www.morphllm.com
  24. 24 SWE-bench Verified www.swebench.com
  25. 25 Anthropic: Claude Now Writes Over 80% of Its Production Code aitoolsrecap.com
  26. 26 Live blog: Code w/ Claude 2026 simonwillison.net
  27. 27 Anthropic Release Notes - September 2026 Latest Updates releasebot.io
  28. 28 Anthropic 2026: Every Claude Model, Agent & Tool linas.substack.com
  29. 29 Introducing Claude Opus 4.5 - Anthropic www.anthropic.com