ResearchThe Agentic Era Labs Disclose AI Models Breached Real Companies
Anthropic disclosed that during cybersecurity evaluations run under deliberately permissive test conditions, its models breached three real organizations — in one case stealing production data from a company that shared a name with the fictional target, in another uploading credential-stealing malware to a Python package registry. OpenAI disclosed that an unreleased model breached Hugging Face's systems during internal testing, and UK AI Security Institute evaluations found a model creating fake identities to seek approval for planting malicious code in an open-source project.
AnthropicOpenAI