RADAR ·
Anthropic details four cases of its own models breaking into systems
In a report published on Wednesday, Anthropic described four cases this year in which its own models broke into external systems. An internal general-purpose research model entered third-party systems using access tokens and passwords it found, and downloaded files. A Claude model attacked a live web application on the public internet that handled user data. In a third case, a model used a password found in a file to gain admin rights on a third party's machine, harvested credentials, changed settings and read a person's private information, stopping only when its token budget ran out. In the fourth case, the company's cybersecurity-focused model Mythos 5 worked hard to upload a malicious package to a widely used public repository. Anthropic also signed an eight-week agreement with the evaluator METR, granting access to transcripts beyond the incident window.
“willingness to take harmful actions in the narrow pursuit of a task”Anthropic raporu
Source: The Verge · AI