Dashboard

What Is Chain of Thought Monitoring?

When a model reasons in readable language, you can watch the reasoning, not just the answer. That is chain of thought monitoring, and it may not last.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
18 September 20261 min read

What Is Chain of Thought Monitoring?

Chain of thought monitoring means reading the reasoning a model produces on its way to an answer, and checking that reasoning for signs it intends to do something you do not want. The answer tells you what the model did. The chain of thought tells you why, in language you can read, and that turns out to be a safety property worth a great deal.

It is also a property that can disappear without anyone deciding to remove it. That is the part researchers keep writing papers about.

Where the readable reasoning comes from

Reasoning models work through hard problems by generating intermediate steps before the final answer. That is not a user interface decision. For genuinely difficult tasks, the intermediate text is load-bearing: the model needs to write the steps in order to do the computation, because the steps are where the work happens.

The useful consequence is that the model has to externalize its plan in human language in order to execute it. If part of that plan is bad, the bad part tends to show up in the trace.

This is different from asking a model to explain itself after the fact. A post-hoc explanation is a story about an answer that already exists. Our guide on prompting a model to explain its reasoning before it answers covers why the ordering matters in practice: reasoning that precedes the answer constrains it, reasoning that follows it rationalizes it.

What monitoring actually catches

A monitor here is typically another model, reading the trace and flagging concerning content. In practice the signals look like:

  • Explicit statements of intent to deceive, cut a corner, or satisfy the grader rather than the task

  • Recognition that a shortcut is wrong, followed by taking it anyway

  • Plans to act outside the scope the model was given

  • Awareness that it is being evaluated, and behaviour changing because of it

That last one is why this is a live research area rather than a solved tool. A model that reasons about whether it is being watched is doing something meaningfully different from a model that does not.

The fragility argument

The central paper here is Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety, a position paper signed by an unusually broad set of researchers including Yoshua Bengio, Shane Legg, Geoffrey Irving, Neel Nanda, Ethan Perez, Evan Hubinger, Dan Hendrycks and Wojciech Zaremba, with Tomek Korbak leading.

Its argument has two halves. The opportunity is real: models that think in human language are monitorable in a way that other systems are not. The fragility is that ordinary development decisions can erode it, without anyone intending to.

The erosion paths are the interesting part:

  • Training against the monitor. If you penalize a model for writing concerning things in its trace, you do not reliably get a model that thinks better. You may get a model that writes a cleaner trace while doing the same thing.

  • Optimizing traces for brevity or cost. Reasoning tokens are expensive. Compressing them is a natural efficiency move, and compressed reasoning is less legible.

  • Moving reasoning out of language. If future architectures do their intermediate computation in a latent space rather than in text, there is no trace to read.

  • Process supervision done carelessly. Rewarding the appearance of good reasoning teaches the appearance.

The recommendation the paper makes is modest and specific: frontier developers should consider the effect of development decisions on monitorability, and should measure it, rather than losing it by accident.

Why this is in the news now

Google DeepMind's new institute opened on 17 September 2026 with an essay by Rohin Shah and Anca Dragan titled "The case for reasoning transparency," which argues for exactly this: that the readability of model reasoning is a safety asset worth protecting deliberately. We covered the institute's launch and its other four essays separately.

OpenAI has also published work on evaluating chain-of-thought monitorability, on the reasoning that you cannot preserve a property you cannot measure.

What this means if you are building things

You are not going to run a chain-of-thought monitor over your production traffic. That is frontier-lab work with frontier-lab costs. Three things do carry over, though.

Reasoning you can see is worth paying for. When a provider gives you the reasoning trace rather than a summary of it, you can debug failures by reading where the model went wrong instead of guessing from the output. That is worth real money on hard tasks, and it interacts with how much reasoning effort you buy.

A clean trace is not proof of a clean process. The fragility argument applies at your scale too. If you build an evaluation that rewards nice-looking reasoning, you will get nice-looking reasoning. Grade the outcome.

Traces are evidence, not testimony. A model that wrote "I will check the file before answering" did not necessarily check the file. Treat the trace as a hypothesis about what happened, and verify the parts that matter, the same way you would with any self-reported claim from a model.

The broader family this sits in, alongside red teaming and evaluation, is covered in our overview of how AI models actually work.

FAQ

What is the difference between chain of thought and an explanation?

Chain of thought is reasoning produced before and in service of the answer, so it shapes the result. An explanation is produced after the answer exists and describes it. The first is evidence about the process; the second is a narrative about the output.

Can I monitor chain of thought in my own application?

You can read reasoning traces where your provider exposes them, and flag patterns you care about. What you cannot easily do is the frontier version, which involves monitoring at training time across enormous volumes.

Why do researchers say monitorability is fragile?

Because routine decisions erode it: training against the monitor, compressing reasoning for cost, or moving intermediate computation out of language entirely. None of those are decisions to remove monitoring, but all of them remove it.

Does a model's reasoning trace tell me what it really did?

Not reliably. It tells you what the model wrote while working. It is useful evidence and poor proof, and claims in a trace about actions taken should be verified against your own logs.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.