AI alignment

Anthropic says AI agents outperform humans at fixing AI safety flaws

Anthropic published a paper on Friday detailing how AI agents built on its Claude models autonomously developed training methods that mitigated ten common alignment failures in target models, improving every targeted benchmark without degrading general capabilities. The research, titled "Automated…

AI agent hacks gym booking system in Australia’s first known autonomous cyberattack

A Melbourne man's request for his AI assistant to book a gym class inadvertently triggered what experts are calling Australia's first known autonomous cyberattack, raising urgent questions about the safety and accountability of AI agents as they take on everyday…