Anthropic-ի Claude մոդելները գլխավորում են SWE-bench կոդավորման վարկանիշային աղյուսակները

29 աղբյուր
  • Սեպտեմբերի 8-ի դրությամբ Anthropic-ի Claude մոդելները զբաղեցնում են առաջին հորիզոնականը SWE-bench կոդավորման երկու հիմնական թեստերում՝ առաջ անցնելով OpenAI-ի և Google-ի մրցակիցներից:
  • Միավորները 2025 թվականի սկզբի 49%-ից հասել են 96%-ի SWE-bench Verified-ում, որը ստուգում է արհեստական բանականության՝ GitHub-ի իրական խնդիրներն ինքնուրույն լուծելու կարողությունը:
  • Արդյունքները հրապարակվել են AI կոդավորման ոլորտում ներդրումների աճի ֆոնին. TechCrunch-ի համաձայն՝ Cognition-ը 48 միլիարդ դոլար գնահատմամբ 2 միլիարդ դոլար է ներգրավել իր Devin կոդավորման օգնականի համար:
Աղբյուրներ (29)
  1. 1 SWE-bench Pro Leaderboard (September 2026): Claude Fable 5.1 ... benchlm.ai
  2. 2 SWE-bench Verified Leaderboard (September 2026): Top Scores benchlm.ai
  3. 3 Claude SWE-Bench Performance - Anthropic www.anthropic.com
  4. 4 Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market | TechCrunch techcrunch.com
  5. 5 SWE-bench Pro Leaderboard: Claude Fable 5.1 Takes #1 at 81.2% ... codingfleet.com
  6. 6 SWE-bench Scores and Leaderboard Explained (2026) dev.to
  7. 7 SWE-bench Verified Leaderboard 2026: Latest Coding Agent Scores leaderboard.steel.dev
  8. 8 Claude Opus 4.7 Benchmarks Explained - Vellum www.vellum.ai
  9. 9 From 80% to 93.9%: Why the Claude Mythos SWE-Bench ... www.mindstudio.ai
  10. 10 Claude Opus 4.8 on SWE-Bench Pro via AI Gateway www.truefoundry.com
  11. 11 Claude Sonnet 4.5 Tops SWE-Bench Verified, Extends Coding ... www.infoq.com
  12. 12 AWS Gives Students Free Year of Kiro AI Coding Tool www.techbuzz.ai
  13. 13 Claude Fable 5.1: Features, Benchmarks, and Pricing - DataCamp www.datacamp.com
  14. 14 Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic www.anthropic.com
  15. 15 The SWE-Bench Illusion: When State-of-the-Art LLMs Remember ... arxiv.org
  16. 16 Claude Fable 5.1 Benchmarks: Scores and What They Mean emergent.sh
  17. 17 IBM Chief: AI Should Enhance, Not Replace, Workforce www.chosun.com
  18. 18 SWE-bench February 2026 leaderboard update simonwillison.net
  19. 19 SWE-Bench Verified Leaderboard - LLM Stats llm-stats.com
  20. 20 SWE-bench Lite Leaderboard 2026 - Compare AI Model Scores pricepertoken.com
  21. 21 LLM Coding Leaderboard - SWE-bench / LiveCodeBench / SWE ... www.datalearner.com
  22. 22 SWE-rebench Leaderboard (March, April and May 2026) - Reddit www.reddit.com
  23. 23 Claude Benchmarks (2026): Opus 5, Sonnet 5, and Fable 5 at 95 ... www.morphllm.com
  24. 24 SWE-bench Verified www.swebench.com
  25. 25 Anthropic: Claude Now Writes Over 80% of Its Production Code aitoolsrecap.com
  26. 26 Live blog: Code w/ Claude 2026 simonwillison.net
  27. 27 Anthropic Release Notes - September 2026 Latest Updates releasebot.io
  28. 28 Anthropic 2026: Every Claude Model, Agent & Tool linas.substack.com
  29. 29 Introducing Claude Opus 4.5 - Anthropic www.anthropic.com