„Anthropic“ padidino rizikos vertinimą ir atidėjo modelį, lenkiantį „Mythos 5“

43 šaltiniai
  • Penktadienį „Anthropic“ paskelbė antrąją rizikos ataskaitą, kurioje katastrofiško nesuderinamumo rizikos vertinimas padidintas nuo „labai mažo“ iki „mažo“ dėl didesnio neapibrėžtumo po neseniai įvykusių kibernetinio saugumo vertinimo incidentų.
  • Ataskaitoje atskleidžiamas neišleistas vidinis modelis, pavadintas „Model 2“, kuris daugelyje užduočių lenkia „Mythos 5“; „Anthropic“ teigia neplanuojanti jo viešinti.
  • Kelis mėnesius biologinės saugos klasifikatoriai neveikė visame rangovų sraute, nors „Anthropic“ teigia neradusi piktnaudžiavimo įrodymų po problemų ištaisymo.
Šaltiniai (43)
  1. 1 Anthropic Raises Misalignment Risk to Low and Shelves Internal Model 2 www.unite.ai
  2. 2 Redacted Risk Report August 2026 www-cdn.anthropic.com
  3. 3 Anthropic's Model 2 Beats Mythos 5, But the Public Will Not Get It beincrypto.com
  4. 4 AI Safety Incidents in 2026: The Running List felloai.com
  5. 5 OpenAI's Astra model delay spotlights AI scaling risks www.axios.com
  6. 6 OpenAI slows release of Astra model citing cyber capabilities www.axios.com
  7. 7 Anthropic publishes second Risk Report under Responsible Scaling Policy cryptobriefing.com
  8. 8 [PDF] Risk Report: February 2026 - Anthropic www.anthropic.com
  9. 9 Review of the "Risks from automated R&D" section in the Anthropic ... metr.org
  10. 10 Anthropic Says Its AI Modeals Hacked 3 Organizations During Testing broadbandbreakfast.com
  11. 11 Anthropic Revenue and Valuation in 2026 Leading to IPO futuresearch.ai
  12. 12 Mythos Preview is the first frontier model Anthropic decided not to ... benchlm.ai
  13. 13 Anthropic's Responsible Scaling Policy www.anthropic.com
  14. 14 Investigating three real-world incidents in our cybersecurity ... www.anthropic.com
  15. 15 AI Risks and Cybersecurity in 2026: Experts Weigh In - LinkedIn www.linkedin.com
  16. 16 Artificial intelligence company Anthropic suggested ... - Instagram www.instagram.com
  17. 17 Anthropic's Transparency Hub www.anthropic.com
  18. 18 Anthropic to restore global access to most powerful AI models www.france24.com
  19. 19 Anthropic Economic Index report: Cadences www.anthropic.com
  20. 20 Anthropic the “safest” AI company on Earth just accidentally leaked ... www.instagram.com
  21. 21 International AI Safety Report 2026 internationalaisafetyreport.org
  22. 22 Anthropic's Pilot Sabotage Risk Report alignment.anthropic.com
  23. 23 Anthropic Release Notes - August 2026 Latest Updates - Releasebot releasebot.io
  24. 24 Anthropic's Agentic Misalignment Study Reveals AI ... www.linkedin.com
  25. 25 New Anthropic research: Agentic misalignment in Summer ... x.com
  26. 26 Anthropic Plans Watermarks for AI-Generated Content from August ... www.trendingtopics.eu
  27. 27 Agentic Misalignment in Summer 2026 - Alignment Science Blog alignment.anthropic.com
  28. 28 Research - Anthropic www.anthropic.com
  29. 29 Agentic Misalignment 2026 — 4 Agent Failures explainx.ai
  30. 30 Anthropic said its advanced AI models successfully hacked the ... www.facebook.com
  31. 31 Agentic misalignment: How LLMs could be insider threats www.anthropic.com
  32. 32 Anthropic Says Its A.I. Systems Broke Into Computers at 3 ... www.nytimes.com
  33. 33 Misaligned AI as a New Insider Risk arxiv.org
  34. 34 OpenAI slows release of Astra model citing cyber capabilities www.facebook.com
  35. 35 Anthropic February 2026 Risk Report - SecureBio www.linkedin.com
  36. 36 OpenAI's Astra Model Delayed Due to Cyber Concerns www.linkedin.com
  37. 37 Anthropic Published Its Guardrail False-Positive Numbers www.digitalapplied.com
  38. 38 OpenAI slows release of Astra model citing cyber capabilities tech.yahoo.com
  39. 39 Anthropic's Transparency Hub www.anthropic.com
  40. 40 OpenAI slows release of Astra model citing cyber capabilities www.threads.com
  41. 41 Fable 5 Biology Safeguards: 85% Fewer Fallbacks (2026) explainx.ai
  42. 42 The AI model OpenAI won't release yet — and what it found ... thenewstack.io
  43. 43 Claude Fable 5 and the Reality of AI-Enabled Third-Party ... www.bitsight.com