Do You Need an LLM for This Task?
Do you need an LLM for this task? Most AI spend goes on calls a classifier handles better. A decision table, the cost arithmetic, and when to pay.
Do you need an LLM for this task? Usually the honest answer is no, or not for all of it. A language model is the most expensive, slowest and least predictable way to turn an input into an output, and it is worth it only when the input is genuinely open-ended. A surprising share of production AI spend goes on calls where the answer is one of four values and a regular expression would have been right every time.
This is a guide to telling the difference before you build, with a decision table, the cost arithmetic, and the three cases where paying for a model is obviously correct.
Start with the shape of the output
Forget the input for a moment and look at what comes back.
Output shape | Cheapest thing that works | When you need more |
|---|---|---|
A fixed value from a known list | Rules, or a trained classifier | Categories overlap or depend on tone |
A number extracted from structured text | Regex or a parser | Format varies wildly across sources |
A number extracted from a document or photo | OCR, then a parser | Layout is unpredictable |
A label plus a confidence score | A classifier or decision model | Decision needs world knowledge |
A ranking of known items | Scoring function or embeddings | Criteria are stated in natural language |
Yes or no on something subjective | A small model | Subtlety genuinely matters |
Prose a human will read | A language model | Always |
Reasoning across several documents | A language model | Always |
The dividing line is not difficulty. It is whether the output space is enumerable. If you can write down every valid answer, you are doing classification, and classification has cheaper tools than generation. If the valid answers are unbounded, you need a model that generates.
The arithmetic nobody runs
Take a support inbox routing 50,000 tickets a month into eight queues.
Through a frontier chat model at roughly 500 input tokens and 10 output tokens per ticket, you are paying frontier rates for 25 million input tokens, plus latency of a second or two per call, plus a parser for when the model answers "Billing (though possibly Technical)."
Through a decision model, you are at a different tier entirely. Cloudflare's Clef and Clef-flash, released on 1 October 2026, price at $0.24 and $0.09 per million input tokens on Workers AI with output tokens not charged, and Clef-flash returns in a median 39 ms, per Cloudflare's announcement. The same 25 million input tokens costs a couple of dollars and returns in a fraction of the time, with the answer constrained to your eight queues by schema rather than by hope.
Through a classifier you train on your own labelled history, the marginal cost is near zero and the latency is a few milliseconds.
Three orders of magnitude separate the ends of that range for the same task. The reason teams land at the expensive end is not analysis, it is that the chat model was already wired up and the prompt took ten minutes to write.
That is a real advantage, and worth being honest about: the expensive option is often the right first version precisely because it is fast to build. The mistake is leaving it there once the volume is known.
Three cases where the task genuinely needs an LLM
The input is prose written by a human who did not follow a format. Any rule you write will meet an input that breaks it. This is what language models are for.
The task needs knowledge you have not encoded. Deciding whether a supplier's email is a price increase requires knowing what a price increase looks like in commercial correspondence. You could build that, or you could use a model that already has it.
The output is language. Summaries, replies, explanations, drafts. There is no cheaper tool, because nothing else produces fluent text.
Everything else deserves a second look.
The hybrid that usually wins
Most real systems should not pick one. They should route.
incoming item
-> cheap deterministic check (does this match a known pattern?)
yes -> done, no model call
no -> classifier or decision model
confident? -> done
not confident? -> language model
still unclear? -> humanEach layer handles what it can and escalates what it cannot. In practice the first layer absorbs more than people expect, because real traffic has a heavy head of repetitive cases and a long thin tail of genuinely novel ones. Paying frontier prices for the head to serve the tail is the specific waste worth fixing.
The escalation only works if each layer can say "I am not sure," which is why having a real probability matters rather than a sentence that sounds hesitant. Setting the cut-off is its own small discipline, covered in how to set a confidence threshold for an AI classifier.
Four questions before you add a model call
Can I enumerate the possible answers? If yes, you are classifying, not generating, and the cheap tools apply.
What is the volume, times the per-call cost, times twelve months? Run the number. The decision looks different at 500 calls a month than at 500,000.
What happens when it is wrong? High-consequence outputs need a confidence signal and a human path, whichever tool produces them. That requirement does not go away by using a better model.
Does this need to be in the request path? Sub-100 ms rules out most generative calls and is comfortable for a small classifier.
Answering these takes twenty minutes and routinely changes the architecture. For the broader decision of which model to reach for once you have established that you need one, frontier model versus cheap model covers that trade, and how to estimate tokens for a task turns the arithmetic above into real numbers for your own workload.
Getting structure out when you do use a model
If you do need a language model but want a constrained answer, structured output forces the response into a schema and removes the parsing problem. It does not remove the cost or latency problem, and it does not give you a calibrated confidence number, but it is the right default whenever a model's answer feeds into code rather than a person.
For the wider picture of putting these pieces together into something that runs, the guide to building an app with AI is the starting point.
FAQ
How do I know if my task is really classification?
Write down every answer the system is allowed to give. If that list is finite and you could hand it to a new colleague as a reference sheet, it is classification.
Is it worth training my own classifier in 2026?
If you have a few thousand labelled examples from your own history and the task is stable, yes. It will be faster and cheaper than any API call and you own it. If you have no labelled data, a decision model or a small model is the faster path.
What about using a cheap model instead of a frontier one?
That is a reasonable middle step and often enough. Just notice it is still generation, with generation's latency and its tendency to answer in prose. If the output is a label, a tool built for labels will beat it on both.
Does this apply to AI features inside my product, or just internal tooling?
Both, and the case is stronger for product features, because they run at user volume and in the request path. A 39 ms classification is invisible to a user; a 1.5 second model call is not.
Will the model approach get cheap enough that this stops mattering?
Prices have fallen steadily and will keep falling, which narrows the gap but does not close it, because the cheaper tools fall too. Latency is the more durable difference: generation has a floor that classification does not.
How did this land?
About the author

Staff Engineer, Platform
Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.


