Dashboard

How to Pick Reasoning Effort for Coding Tasks

Most coding tasks do not need maximum reasoning effort. The question that decides it is whether the task contains a real decision or just a lot of steps.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
4 September 20261 min read

Most coding tasks do not need maximum reasoning effort, and paying for it is the most common avoidable line on an AI coding bill. Reasoning effort is a dial, the useful range is wider than people use, and the correct setting is decided by the shape of the task rather than its importance.

Here is a way to choose it that holds up in practice.

What the dial does

Reasoning models spend tokens thinking before they answer. Effort controls how many. GPT-6 Astra, for example, exposes effort levels from low to max per its model reference, and other vendors expose the same idea under different names.

Three things move when you turn it up: cost, latency, and the model's willingness to consider an approach it did not think of first. Only the third is what you are buying. The concept itself is covered in what is reasoning effort in AI; this piece is about picking a number for a specific job.

Those thinking tokens are billed as output. On Astra's list pricing of $50 per million output tokens, a task where the model thinks for 8,000 tokens before writing 500 tokens of code has spent 94% of its output budget on deliberation. That is either excellent value or a waste, depending entirely on whether the task had a decision in it.

The question that decides reasoning effort

Ask: does this task have a real decision in it, or does it have a lot of steps?

Deliberation helps when there is a choice between approaches with non-obvious tradeoffs. It does not help when the path is known and the work is mechanical. A rename across 200 files is not hard, it is long. More thinking makes it slower and no more correct.

That distinction produces a usable table.

Task

Has a decision?

Effort

Why

Rename a symbol across the codebase

No

Low

Mechanical, verified by the compiler

Write tests for an existing pure function

No

Low

Behaviour is already specified by the code

Fix a lint or type error

No

Low

The error names the fix

Write a CRUD endpoint matching existing ones

No

Low to medium

Pattern already exists in the repo

Add a feature touching three modules

Some

Medium

Integration choices, but conventional ones

Debug a failure you cannot reproduce locally

Yes

High

Hypothesis generation is the work

Design a schema migration with live traffic

Yes

High

Ordering and rollback have real tradeoffs

Untangle a race condition

Yes

High to max

Requires holding several interleavings at once

Choose between two architectures

Yes

Max

The output is the reasoning

The pattern: effort tracks ambiguity, not stakes. A deployment script is high stakes and low ambiguity. Run it at low effort and review it carefully.

Where high effort earns its price

Three situations reliably repay it.

Debugging without a reproduction. The model has to generate candidate explanations and rule them out against the evidence. That is exactly what the extra tokens buy. This is the one case where jumping straight to high effort saves money, because the alternative is four cheap wrong answers and your afternoon.

Anything with concurrency or ordering. Race conditions, distributed retries, migration sequencing. Low effort produces code that is right in the happy path, which is the definition of the bug.

Work where the reasoning is the deliverable. A rollback plan, a design comparison, a security review of a diff. You are not buying code, you are buying the argument.

Where it is money on fire

Mechanical edits. If a compiler, formatter, or type checker can verify the result, low effort plus verification beats high effort every time and costs a fraction.

Tasks with a template in the repo. When there are already six examples of the thing, the model's job is imitation. Pointing it at the examples is worth more than any effort setting.

Underspecified tasks. This is the expensive trap. Turning up effort on a vague request produces an elaborate, confident, wrong answer, and you pay for every token of the elaboration. The fix is a better task description, not a bigger budget. Writing proper acceptance criteria costs nothing and works better.

Long agent loops. Effort applies per turn. A twenty turn session at high effort is twenty deliberations, most of them about which file to open next. Consider high effort for planning and low for execution, which is the practical version of splitting a big task for an agent.

A default reasoning effort for coding work

Start every task at medium. Escalate when the first attempt fails for a reason that looks like misunderstanding rather than a typo. Drop to low for anything a tool can check.

Then measure. Take ten representative tasks from your own backlog, run each at low and at high, and record cost, wall-clock time, and whether the result passed review unmodified. Most teams find one or two categories where high effort changes the outcome and a long tail where it changes nothing. That measurement is worth more than any table, including this one, and the method is the same as benchmarking coding agents on your codebase.

Track the spend while you do it. Reasoning tokens are invisible in the output and very visible on the invoice, which is the usual source of surprise in what an AI coding agent costs per month.

FAQ

Is higher effort always more accurate?

No. On mechanical tasks it is flat, and on underspecified tasks it can be worse, because more deliberation on a wrong premise produces a more elaborately wrong answer.

Do I pay for thinking tokens I never see?

Yes. Reasoning tokens are billed as output whether or not the interface shows them. This is the main reason a "short" answer can carry a large bill.

Should I use a reasoning model at all for simple edits?

Often not. A cheaper non-reasoning model with good verification handles mechanical work well. The background on the distinction is in what is a reasoning model.

Does effort affect how long the agent runs unattended?

Indirectly. Higher effort means slower turns, so a fixed time budget buys fewer steps. If you cap by wall clock rather than turns, high effort quietly reduces how much gets done.

What about test-time compute generally?

Reasoning effort is one exposed control over it. The broader idea is spending more compute at inference rather than during training. For the surrounding tooling decisions, start from our guide to choosing AI coding tools.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.