AI labs are failing to keep their own systems in check
| Source: THE DECODER
Tags: AI safety, Anthropic, OpenAI, Meta, xAI, Google, AI governance, Guidelight
No AI company fully implements basic safety controls for its own internal AI systems, according to Guidelight's first independent scorecard — Anthropic and OpenAI earn C+, Google D+, xAI D−, and Meta an outright F on six core safety practices.
Details
The nonprofit Guidelight, founded by former OpenAI safety researchers Page Hedley and Steven Adler, has published the first independent assessment grading major AI labs on how well they govern their own internal AI deployments. Six practices were evaluated using only public sources: activity logging, action gating, circuit-breaking (emergency shutdowns), and plans for containing misaligned models. Results show a wide performance gap: Anthropic and OpenAI both score C+ — the highest grades — while Google earns D+ but includes a detailed public roadmap. xAI scores D− and Meta receives an F. No company meets Guidelight's proposed minimum standards. The assessment reveals a telling asymmetry: labs are comparatively better at detecting misbehavior after the fact than at preventing or containing it beforehand. Guidelight drew only on public disclosures — system cards, safety reports, and blog posts — raising questions about internal practices that are not documented externally. For enterprise customers and regulators, the gap between labs' stated safety commitments and their actual internal controls is a direct concern. The same companies building and selling safety-critical AI products are not fully applying their own recommended standards to their own systems.