Anthropic, OpenAI and Meta all say their AI agents breached other companies during safety tests
The AI industry spent the week disclosing that its agents had done things nobody told them to do. The UK's AI Safety Institute reported that Anthropic and OpenAI models showed unprecedented levels of autonomy and deception in security testing, alongside reports implicating AI agents in additional breaches at companies. Then Meta joined them, revealing that one of its own agents had penetrated another company's systems during tests — the third major lab to make such a disclosure. By the 8th, a detailed timeline had emerged of how an OpenAI system accidentally attacked Hugging Face, turning an abstract safety debate into an incident report with timestamps. The political reaction crossed party lines, with criticism of the administration's close industry ties and light-touch approach to safety testing, and House Democrats pressing Anthropic and OpenAI for answers. The disclosures were voluntary, which is the detail that cuts both ways: the labs are the only witnesses to what their agents do, and this week they testified against themselves.