Dashboard

How Much It Costs to Run an AI Agent 24/7

How much does it cost to run an AI agent 24/7? Build it from three meters: wake frequency, tokens per wake and tool calls. Idle time usually wins.

Manuele Estivo
Manuele Estivo
Growth & SEO Lead
30 September 20261 min read

The cost to run an AI agent 24/7 is driven less by the work it does than by how often it wakes up to check whether there is work. Most people estimate an always-on agent by pricing the tasks they expect it to complete, which is the wrong denominator. A useful model has three meters: how often it wakes, how many tokens each wake costs, and how many tool calls each wake triggers. Get those three and the monthly number falls out. Do the arithmetic honestly and idle checking often costs more than the actual output.

The three meters that decide your bill

Everything else is detail. Write these down before you price anything:

  1. Wake frequency. How many times in 24 hours does the agent become active? A one-minute polling loop is 1,440 wakes a day. Hourly is 24. That is a sixtyfold difference in everything downstream.

  2. Tokens per wake. Critically, this includes the context you resend every time: the goal, the state, the history, the tool definitions. This is the number people forget, and it does not shrink just because nothing happened.

  3. Tool calls per wake. Each call adds its own round trip, its own returned payload, and often its own third-party charge.

Notice that only the third meter has anything to do with useful work. The first two are charged whether or not there was anything to do, which is the structural difference between an always-on agent and a task you run on demand.

What it costs to run an AI agent 24/7, in arithmetic

Take a concrete shape. An agent watches an inbox, and you want it responsive, so it checks every minute. Its standing context is modest: the goal, a summary of state, and the tool definitions, call it 4,000 tokens. Substitute your provider's current input price per million tokens, which we will call P.

Wake interval

Wakes per 30 days

Idle input tokens per month

Cost at P per million

Every minute

43,200

172.8 million

172.8 x P

Every 5 minutes

8,640

34.6 million

34.6 x P

Every 15 minutes

2,880

11.5 million

11.5 x P

Hourly

720

2.9 million

2.9 x P

That table is the whole lesson. Moving from one minute to fifteen cuts standing cost roughly fifteenfold and, for an inbox, changes nothing a human would notice. Meanwhile the emails it actually drafts might be forty a day, a rounding error against 172.8 million tokens of looking.

We are deliberately not inserting a vendor price here, because per-token pricing moves monthly and a number in a blog post ages badly. Put your current rate in and the ranking of the rows will not change. If you need help sizing the per-task side, how to estimate tokens for an AI task has the method.

Why the standing context is the expensive part

An always-on agent has to be reminded who it is on every wake. That reminder grows, because the natural fix for an agent losing the thread is to give it more history, and history is charged by the token every single time.

Two compounding effects make this worse than linear. Long inputs cost more per unit of usefulness, for the reasons in why long context costs more. And an agent running for days accumulates state that nobody prunes, so the 4,000 tokens in the table above is a starting figure, not a steady one.

Practical consequence: the single highest-leverage cost control on an always-on agent is a hard cap on standing context, enforced in code, with old state summarised or dropped rather than carried.

This is a different question from app hosting

Worth separating two things people conflate. What it costs to run an AI-built app is mostly hosting, database and bandwidth: costs that scale with your users. An always-on agent's cost scales with the clock, independent of whether anyone used your product that day. A quiet weekend reduces the first bill and does nothing to the second.

It is also a different question from a per-seat subscription. What an AI coding agent costs per month is a licence you buy. This is consumption you cause.

Six ways to cut the number without losing responsiveness

  • Replace polling with events. A webhook that fires on the thing you care about turns 43,200 wakes into however many events actually happened. This is the single biggest win available and it is usually an afternoon of work.

  • Widen the interval until someone complains. Most teams pick one minute because it felt safe, never because anyone measured the tolerable delay.

  • Use a two-stage check. A cheap, small model decides whether anything is worth escalating, and only then does the expensive model wake up properly.

  • Cap and summarise standing context on a schedule, so the per-wake cost stops growing.

  • Use caching for the parts of your context that never change, such as tool definitions and the goal.

  • Pick the cheaper model for the loop. Vendors keep shipping near-flagship models at a fraction of flagship prices, and a watcher does not need the flagship.

On that last point, OpenAI's GPT-6.1 Sol was announced at roughly a fifth of Astra's token cost while claiming close to Astra-level performance, per BGR's DevDay coverage. Conversely, the new premium speed tier costs about six times standard API pricing, which is precisely the wrong thing to put in a background loop. We worked through that trade in our piece on OpenAI Ultrafast.

The number your vendor has not given you

If you are using a bundled agent rather than building on an API, the cost question changes shape and gets harder to answer. When OpenAI announced its always-on Dots, it did not publish the usage allowance for intensive work, the price of additional agents, or what higher workload capacity costs, as VentureBeat noted at launch.

So the honest answer for bundled agents is that you cannot yet model the cost, only the failure. Run one on something non-critical and watch for the point where it stops. That boundary is your real capacity, and knowing it is worth more than an estimate.

The hidden multiplier: retries and tool payloads

Two line items routinely double a careful estimate. The first is retries. An agent that fails a tool call and tries again pays for the whole context twice, and a flaky third-party API turns a predictable loop into an unpredictable one. Budget a retry factor rather than assuming the happy path.

The second is what tools return. A search tool that hands back ten full pages puts all ten into the next request. The agent did not choose that, your tool definition did, and truncating tool output at the boundary is often a bigger saving than any change to the model or the prompt.

If you are billing a client for this

An always-on agent is the worst possible thing to put inside a flat monthly fee, because its cost is set by a configuration choice that the client can ask you to change. If a client asks for faster checking, that is a price change, not a preference. Metering it separately, or attaching the polling interval to the price, keeps a configuration change from quietly eating your margin. The wider strategy sits in our AI monetization guide.

FAQ

How much does it cost to run an AI agent 24/7?

It depends almost entirely on wake frequency and the context resent each wake, not on how much work it completes. A one-minute polling loop with a 4,000-token standing context spends roughly 173 million input tokens a month before doing anything useful. Widening to fifteen minutes cuts that about fifteenfold.

Is it cheaper to keep an agent running or start it on demand?

On demand is almost always cheaper, and for most jobs it is also fast enough. Continuous running earns its cost only when the delay between something happening and the agent noticing genuinely matters, and when events cannot be pushed to you.

Does an idle AI agent cost anything?

If it polls, yes, and this surprises people. Each check resends its whole standing context and is charged in full even when the answer is that nothing changed. An event-driven agent genuinely costs close to nothing while idle.

What is the fastest way to cut an always-on agent's bill?

Switch from polling to events, then cap the standing context. Those two changes usually take a day and routinely cut an order of magnitude, which no amount of prompt tuning will match.

How did this land?

About the author

Manuele Estivo
Manuele Estivo

Growth & SEO Lead

Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.