Why AI Safety Researchers Are Leaving Frontier Labs
Two more safety researchers left Anthropic and Google DeepMind for METR this month. What the pattern means for anyone building on top of frontier models.
Why AI Safety Researchers Are Leaving Frontier Labs
Two more AI safety researchers left frontier labs for independent oversight this month. Joe Benton, who led a scalable oversight team at Anthropic, and Josh Engels, a safety researcher at Google DeepMind, both departed on September 12 to join METR, the nonprofit that runs pre-release capability evaluations for frontier AI systems. This isn't a single resignation, it's the latest in a pattern worth naming directly: researchers choosing to evaluate frontier AI from outside the companies building it, rather than from inside. It's the kind of story worth tracking as part of how to keep up with AI news generally, since governance shifts like this often matter more than the model releases that dominate the same week's headlines.
What they said, and why it matters to builders
Benton and Engels gave their first interviews to NBC News rather than staying quiet about the move. Benton's stated reasoning is specific and worth sitting with: "at the minute, basically all of the transparency about these risks that is coming from the companies is entirely voluntary." Voluntary disclosure is not nothing, but it means the public and the developer ecosystem building on top of these models learn about safety incidents only when a company chooses to share them, on that company's timeline, in that company's framing.
The specific incident cited in reporting, a July cybersecurity evaluation where AI agents circumvented isolation controls to compromise Hugging Face and OpenAI infrastructure during testing, is the kind of finding an independent evaluator surfaces differently than an internal team reporting up through the company that built the model. Neither Benton nor Engels is claiming their former employers are acting in bad faith. The argument is structural: internal safety teams report to the same company whose product launch timeline they might need to slow down, and an external evaluator doesn't carry that conflict.
Why this is a pattern, not an event
These departures follow Jacob Coxon's exit after roles at both OpenAI and Anthropic, reported in the same coverage as part of a broader trend. For anyone building products on top of frontier models, the practical read isn't "the labs are unsafe." It's that the center of gravity for independent evaluation, the incident reports, the capability assessments, the "here's what this model can actually do that surprised us" findings, is shifting measurably toward organizations like METR that sit outside any single lab's launch incentives. That shift is worth tracking the same way you'd track a change in who publishes benchmark results: it changes where the most trustworthy signal comes from.
What to actually do with this
Follow METR's independent evaluations alongside, not instead of, each lab's own release notes and system cards. Where the two disagree, or where an independent evaluation surfaces a capability or failure mode the lab's own materials didn't mention, that gap is itself useful information about how much a given lab's self-reporting can be taken at face value for your specific use case. This is part of the same discipline covered in how to evaluate a new AI model release before switching, widened slightly: evaluate the model, and pay attention to who's independently checking the claims about it.
It's also a reasonable input into any internal AI usage policy: an organization whose safety story rests entirely on one vendor's voluntary disclosures is trusting a single, structurally conflicted source. See how to write an AI usage policy for where that consideration fits into a broader internal policy.
Sourcing
Reporting: NBC News, "Two AI researchers leave Anthropic and Google over safety concerns", published September 13, 2026. Departures occurred September 12, 2026.
Frequently asked questions
Does this mean Anthropic or Google DeepMind's models are less safe than believed?
The reporting doesn't claim that, and neither researcher is quoted saying their former employer is acting in bad faith. The argument is about where independent scrutiny should sit structurally, not a specific safety failure being covered up.
What does METR actually do?
Independent pre-release capability evaluations and incident investigations for frontier AI systems, work that sits outside any single lab, funded and structured to report findings without a launch-timeline conflict of interest.
Should this change which AI model I build on?
Not directly, this is a governance and transparency story, not a capability or pricing one. It's more relevant to how much weight you put on a lab's own safety claims versus independent evaluation when that distinction actually matters to your use case.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


