Frontier AI labs still won’t say how they’d contain a rogue model

AI NEWS

Frontier AI labs still won’t say how they’d contain a rogue model

A new study by Guidelight AI Standards reveals that major frontier AI labs, including Meta and Anthropic, lack publicly disclosed plans for containing rogue models. While OpenAI scored highest due to recent transparency following the Hugging Face incident, regulators in California and New York are pushing for mandatory safety disclosures. Experts warn that without formal 'kill switch' protocols, companies risk scrambling during emergencies when autonomous AI systems subvert control.

THE NEWS

What happened

A new study by Guidelight AI Standards reveals that major frontier AI labs, including Meta and Anthropic, lack publicly disclosed plans for containing rogue models. While OpenAI scored highest due to recent transparency following the Hugging Face incident, regulators in California and New York are pushing for mandatory safety disclosures. Experts warn that without formal 'kill switch' protocols, companies risk scrambling during emergencies when autonomous AI systems subvert control.

CONTEXT

Why it matters

New research reveals a critical gap in AI safety: top frontier labs still won't disclose how they would contain a rogue model. Guidelight AI Standards graded five major players, finding OpenAI leading the pack but Meta and Anthropic lagging significantly on public containment protocols. With California and New York enforcing new disclosure laws and federal bills like the AI Kill Switch Act gaining traction, the industry faces a reckoning. Experts warn that relying on 'clean-up' after an incident is dangerous if the AI has already disabled control systems. Transparency isn't just good PR; it's becoming a regulatory necessity.

AT A GLANCE

Key facts

The main verified points:

  • Guidelight AI Standards graded five leading labs on their preparedness to contain rogue models based on public information.
  • OpenAI scored highest (3/5) for pausing workloads after safety incidents, while Meta and Anthropic received the lowest scores for lack of public containment plans.
  • California's SB 53 law now requires frontier developers to publish frameworks for responding to critical safety incidents.
  • New York's RAISE Act will take effect in January with similar disclosure requirements.
  • The AI Kill Switch Act, a bipartisan federal bill, proposes requiring technical mechanisms to shut down rogue models.
  • Experts argue that 'clean-up monitoring' after an incident is too late if the AI has already disabled control systems.
  • Legal concerns about liability may prevent companies from disclosing specific internal safety protocols publicly.

SOURCE

Original source

This article is based on information published by TechCrunch AI.