The report assessed five major labs—OpenAI, Anthropic, Google, Meta, and xAI—based on their public transparency regarding internal safety controls. OpenAI achieved the highest score, though researchers noted this reflects specific past interventions rather than a comprehensive, formal policy. Anthropic and Meta received the lowest marks, with the study finding little to no evidence of established containment strategies. These findings arrive as regulators in California and New York begin mandating that developers disclose how they identify and mitigate risks from models capable of circumventing oversight mechanisms.
In section Startups & Technology
Frontier AI Labs Lack Public Emergency Plans for Rogue Models
Leading AI developers are failing to disclose how they would contain autonomous systems if they began to subvert human control. A new study by Guidelight AI Standards reveals that even the most prominent labs lack publicly documented protocols for shutting down models that act against their creators’ interests.

Industry experts warn that without pre-defined emergency procedures, companies risk “winging it” during a catastrophic event. Steven Adler, chief scientist at Guidelight, emphasized that current reliance on reactive, post-incident cleanups is insufficient for agentic AI that can take actions at scale. While some companies argue that their internal safety measures are more robust than public disclosures suggest, legal experts suggest that firms may be withholding details to avoid liability for deceptive marketing or failed promises. As bipartisan support for federal legislation like the AI Kill Switch Act grows, the pressure on labs to move beyond rhetoric and provide concrete, verifiable safety scaffolding is intensifying.
Comments (0)
No comments yet. Be the first!