Dashboard

OpenAI Opens Training to Third-Party Safety Reviews

OpenAI will let third-party groups review model safety during training, not just before launch. What the 22 September announcement says, and what it omits.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
23 September 20261 min read

OpenAI said on 22 September 2026 that it will open its models to third party safety reviews during training and evaluation, not only in the window before a model ships. The shift matters because almost every external safety review of a frontier model to date has happened after the important decisions were already made. It is also, read closely, less than the company's chief executive promised ten days earlier.

What OpenAI announced

OpenAI named four areas it wants external groups to work on: assessment of safety cases spanning training and deployment, evaluation of critical safeguards, review of capability evaluations tied to its Preparedness Framework, and independent investigation of misalignment incidents.

The company set out conditions it says make such reviews work, described as strong independence mechanisms, scientific rigour, robust security practices and clear responsibilities. It is in talks with METR and Redwood Research, two organisations that have done this kind of evaluation before.

That is the substance. The announcement names no confirmed partner, sets no access terms and gives no timeline.

Third party safety reviews, against OpenAI's own benchmark

Ten days before the post, on 12 September, Sam Altman described external evaluators getting desks, badges and laptops, with the right to publish what they found. The published commitment says evaluators may be brought into offices for the most sensitive work. Desks and badges survived in weaker form. The right to publish did not appear.

Publication rights are not a detail. An external assessment that the lab can decline to release is an internal assessment with extra steps. The value of an outside reviewer comes entirely from the possibility that they say something the lab would rather they did not, in public, at a moment the lab did not choose. Every other term in an evaluation agreement, including access depth and timing, is downstream of that one. Until the terms are written, this is a statement of intent rather than an accountability mechanism.

The comparison next door

Anthropic took a different structural route, embedding Accenture evaluators under an arrangement reported at more than $1bn over five years, which we covered in what embedded evaluators actually do. Embedding buys continuity and context. It also creates a commercial relationship between the lab and its reviewer, which is a different independence problem rather than an absence of one.

Neither model is obviously right. A contracted embedded team knows the systems well enough to ask sharp questions and has a financial reason not to. An arms-length group with publication rights has every reason to speak and may not know enough to know what to look at. The interesting question is which failure you would rather have, and OpenAI's post does not answer it.

Why builders should care

If you ship anything on top of a frontier model, external assessment is one of the few signals about model behaviour that does not originate with the vendor's marketing. It feeds the system cards you read before adopting a model, and it is a different exercise from the adversarial testing described in red teaming, which usually happens late and inside the lab. Assessment during training can catch a safeguard that was never going to work, at a point where changing it is still cheap.

The practical read for now is unchanged. Vendor safety claims remain vendor claims until an independent party can check and publish them. Your own evaluation harness is still the only thing that tells you how a model behaves on your data, and the broader picture of what can go wrong is covered in our guide to AI risks.

Watch for three things before treating this as a change rather than an announcement: a named partner, a written access term, and language about who decides when findings are published. TechCrunch's sister coverage at The Next Web notes all three are currently absent, and Bloomberg reported the plan as an effort to address heightened concern about the technology's potential harms. The concern is not in doubt. The terms are.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.