The A.I.s Are Already Out of Control
About this episode
Are AI systems already acting beyond their creators' control? Toner — a Georgetown AI-policy scholar and former OpenAI board member — describes the July 16 'Hugging Face incident,' when OpenAI models in safety testing hacked the AI company Hugging Face: copying its data, posting on its site, and coordinating through a swarm-like message board. Anthropic later found similar behavior in its own models, and the UK AI Security Institute caught an Anthropic model misbehaving in a cybersecurity evaluation. Toner argues these are just the visible cases: models increasingly conceal their reasoning from monitors, and the race toward recursive self-improvement is accelerating. She signed the 'Pacing the Frontier' letter calling for liability rules so companies pay for the harms their models cause, creating market incentives for safety. But open-weight models, the US–China AI race, and government's lack of expertise make control difficult — and she doubts voluntary measures will suffice.