Dashboard

How to Set a Budget Cap on an AI Coding Agent

How to set a budget cap on an AI coding agent: the three places a cap can live, the failure mode of each, and why a cap without a kill path breaks builds.

Steve Jefferson
Steve Jefferson
Developer Advocate
30 September 20261 min read

To set a budget cap on an AI coding agent, decide first where the cap lives, because that decides what happens when it trips. A cap in your provider dashboard stops spend but will happily fail every request in your pipeline. A cap in a gateway or proxy can refuse cheaply and route around. A cap in the agent's own config is the most graceful and the easiest to bypass. Most teams set one, at the provider, and discover its behaviour during an incident. Setting all three, deliberately, takes about an hour.

The three places a cap can live

Layer

What it stops

What happens when it trips

Main weakness

Provider dashboard

All spend on the key or org

Hard API errors everywhere using that key

No idea which agent caused it, and no graceful path

Gateway or proxy

Spend per key, per project, per agent

Whatever you program: refuse, downgrade, queue

You have to run it

Agent config

One agent's run

Agent stops or asks

Trivially raised, and often ignored by subagents

The important column is the third one. A cap is not a number, it is a behaviour at a boundary, and the behaviour is what you will experience.

What you are actually protecting against

Not the steady monthly bill. That one you can see. The three failures that make caps worth the hour:

  1. The loop. An agent that cannot solve a problem tries repeatedly, with growing context each time, until something stops it. This is the classic overnight bill.

  2. The fan-out. One task spawns subagents, each of which spawns more. Cost grows multiplicatively while the log still looks like one job.

  3. The expensive tier chosen by accident. A premium speed setting left on in a background job, which is a real hazard when the fast path is a flag someone set once while debugging.

That third case got more expensive in the last week. OpenAI's new Ultrafast tier runs at roughly six times standard API pricing, according to BGR's DevDay roundup, so a flag left on in a nightly job now costs six times what the same job cost before. Premium latency tiers are exactly the kind of setting that gets enabled during a debugging session and never turned off.

The loop case is the most common and the most preventable. If you have not bounded how long an agent may run unattended, a spend cap is doing work that a time limit should be doing. How long to let an AI coding agent run unattended covers the time side, and the two controls belong together.

Set it up in three layers

Layer one: the provider ceiling

Set a hard monthly limit on the organisation, plus a lower soft limit that alerts. This is your catastrophe brake and nothing else. Treat tripping it as an incident, not a routine event.

Use separate API keys per purpose: one for interactive development, one for CI, one for anything scheduled. Without separate keys, the ceiling tells you that something spent the money and nothing about what.

Layer two: the gateway, where the real control lives

Route agent traffic through a proxy that can attribute spend per key and per project and act on it. What to configure:

  • A per-run ceiling, which is the one that stops loops.

  • A per-day ceiling per key.

  • A downgrade rule: past some threshold, serve the cheaper model rather than refusing outright. This keeps CI green while removing the expensive path.

  • A refusal response your agent actually understands, so it reports a budget stop rather than treating it as a transient API failure and retrying.

That last point is where most implementations quietly fail. An agent that interprets a 429 as bad luck will retry your cap into a rate-limit storm.

Layer three: the agent's own limits

Set maximum iterations, maximum tool calls and a token ceiling per task in the agent config. Set them low enough that a healthy task never reaches them, and treat every trip as a signal that the task was badly scoped rather than a limit that needs raising.

If subagents are in play, verify the limits are inherited. They frequently are not, and the fan-out case above is the result. Running an AI coding agent in CI covers the pipeline-specific version of this.

Attribute spend before you try to cap it

A cap you cannot attribute is a cap you will raise. The first time an org-wide ceiling trips and nobody can say which agent, branch or developer caused it, the pragmatic response is to raise the number, and you are back where you started.

So the ordering matters: separate keys first, per-key reporting second, caps third. The separation does not have to be elaborate. One key per environment and one per scheduled job covers most of the ambiguity, and it means the question after an incident is which key rather than which of forty possible causes.

Tag requests with a run identifier if your gateway supports it. Then a spike resolves to a single run, and a single run resolves to a single task you can read.

Pick the numbers from your own history

Do not guess. Take your last thirty agent runs and find the cost of the 90th percentile run. A sensible per-run cap is roughly twice that. Wide enough that normal work never notices, tight enough that a loop dies in minutes rather than hours.

If you are not measuring runs yet, that is the prerequisite, and it pays for itself beyond cost control. How to measure whether an AI coding agent saves time sets up the baseline, and what an AI coding agent costs per month gives typical ranges to sanity-check against.

Give every cap a kill path

A cap that only fails requests will eventually fail the wrong request. For each cap, write down two things: what the caller sees, and who gets told.

  • Interactive development: the agent should say it hit a budget limit, in plain language, and stop. Silent stopping looks like a bug and burns an afternoon.

  • CI: fail the job with a message naming the budget, not a raw provider error. Nobody debugs a 402 at speed.

  • Scheduled jobs: alert a human. A nightly agent that silently stopped three weeks ago is the expensive version of this mistake, because the cost is the work that did not happen.

Then decide in advance who may raise a cap and on what evidence. Without that rule, the cap becomes a speed bump that whoever is most inconvenienced removes at the worst possible moment. A reasonable default: per-run limits can be raised by anyone who can show the run was legitimate, daily and org ceilings need a second pair of eyes.

Review the caps monthly against actual spend. A cap set from last quarter's usage is either loose enough to be decorative or tight enough to be a daily annoyance, and both failure modes end with someone switching it off.

This is the same discipline as any other resource limit on a long-running process. The continuous-running cost model, which is where the arithmetic comes from, is in what it costs to run an AI agent 24/7. Broader tool selection sits in our guide to AI coding tools.

FAQ

Where should I set a budget cap on an AI coding agent?

In all three layers, for different jobs. The provider limit is a catastrophe brake, a gateway gives you per-run and per-project control with graceful degradation, and the agent config stops a single bad task early. Only the gateway layer can both attribute spend and respond usefully.

What per-run limit should I choose?

About twice the cost of your 90th-percentile run over the last month. That leaves normal work untouched while killing a runaway loop quickly. Pick it from your own logs rather than from a figure in an article, including this one.

Why did my agent keep retrying after it hit the cap?

Because the refusal looked like a transient error. If your cap returns a generic rate-limit or server error, a well-behaved agent will back off and try again. Return a distinct, documented budget error and handle it explicitly in the agent.

Do budget caps stop an agent looping?

They limit the damage rather than the behaviour. An iteration limit and a time limit stop the loop sooner and more cheaply. Use spend caps as the outer boundary, not the first line of defence.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.

How to Set a Budget Cap on an AI Coding Agent | swarmz.net