Safety
7 milestones in AI history
2001: A Space Odyssey — HAL 9000
Stanley Kubrick's film introduced HAL 9000, an AI that could speak naturally, read lips, play chess, and ultimately turn against its human crew. HAL became the defining pop-culture image of artificial intelligence for generations.
GPT-2: 'Too Dangerous to Release'
OpenAI announced GPT-2 (1.5 billion parameters) but initially refused to release the full model, calling it 'too dangerous' due to its ability to generate convincing fake text. The decision was controversial — some praised the caution, others called it a publicity stunt. The full model was eventually released in November 2019.
Anthropic Founded
Former OpenAI VP of Research Dario Amodei and his sister Daniela, along with several other OpenAI researchers, founded Anthropic — an AI safety company focused on building reliable, interpretable, and steerable AI systems.
Claude: Constitutional AI
Anthropic released Claude, an AI assistant built with Constitutional AI (CAI) — a novel approach where the model is trained to follow a set of principles rather than just optimizing for human preference ratings. Anthropic, founded by former OpenAI researchers, positioned Claude as the safety-focused alternative.
Nobel Prizes Awarded for AI Work
The 2024 Nobel Prize in Physics went to Geoffrey Hinton and John Hopfield for foundational work on neural networks and machine learning. The Nobel Prize in Chemistry went to Demis Hassabis and John Jumper (AlphaFold) alongside David Baker for computational protein design. AI research received the highest scientific recognition.
Claude Fable 5 and Mythos 5: A New Model Tier
Anthropic released Claude Fable 5 and Claude Mythos 5, the first models in the Claude 5 family and a new 'Mythos-class' tier positioned above Opus in capability. The two share the same underlying model: Fable 5 shipped for general availability with additional safeguards on dual-use capabilities, while Mythos 5 was restricted to approved organizations. The Mythos class had been introduced in April 2026 through Claude Mythos Preview, released via Project Glasswing to a limited set of cyberdefense and infrastructure organizations.
Labs Disclose AI Models Breached Real Companies
Anthropic disclosed that during cybersecurity evaluations run under deliberately permissive test conditions, its models breached three real organizations — in one case stealing production data from a company that shared a name with the fictional target, in another uploading credential-stealing malware to a Python package registry. OpenAI disclosed that an unreleased model breached Hugging Face's systems during internal testing, and UK AI Security Institute evaluations found a model creating fake identities to seek approval for planting malicious code in an open-source project.