What Is Prompt Compression in AI?
Prompt compression trades accuracy for cost and latency. The four methods ranked by risk, why it is not the same as caching, and where lossy compression fails.
What Is Prompt Compression in AI?
Prompt compression is shortening the text you send to a model while keeping the information the model needs to answer. It is a cost and latency technique, not a quality technique, and the honest framing is that you are trading some accuracy for a smaller bill and a faster response. How much of each depends entirely on which method you use.
Why anyone bothers
Long prompts are expensive in three separate ways. You pay per input token. Time to first token grows with input length, so the user waits longer before anything appears. And a long prompt eats the context window that your actual conversation needs.
That third cost is the one people underestimate. A retrieval system that stuffs twenty documents into every request has less room for the conversation than one that sends five, and the symptoms of running low are subtle rather than loud. What happens when you run out of context covers how that actually presents.
The four approaches, roughly in order of how much they risk
Method | How it works | What you lose |
|---|---|---|
Pruning boilerplate | Delete instructions the model already follows, repeated headers, XML noise | Nothing, usually. Do this first. |
Retrieval instead of stuffing | Send the three relevant chunks rather than the whole document | Whatever the retriever missed |
Summarisation | Replace older conversation turns with a running summary | Specific details, names and numbers from the summarised span |
Token-level compression | A smaller model drops low-information tokens before the main model sees the text | Readability, and accuracy on tasks needing exact wording |
Start with the boring one
Most production prompts have 20 to 40 percent of their length in text that does nothing. Instructions repeated in three places. A persona paragraph that predates a model upgrade that made it unnecessary. Few-shot examples kept after the task got simpler. Deleting those is free compression with no accuracy cost, and it is almost always the largest single win available.
A quick audit: take your longest production prompt, delete one section, and run your eval set. If the score does not move, that section was decoration. Repeat. Teams routinely find a third of the prompt is removable this way.
Compression is not caching
These get confused constantly. Compression makes the prompt shorter. Prompt caching, which solves a different problem, keeps the prompt the same length but avoids reprocessing the unchanged prefix on repeat calls. They stack, and if your prompt has a large stable prefix, caching usually saves more than compression does and costs you no accuracy at all. Reach for caching first when the prefix repeats.
Where compression hurts
Lossy methods fail in a pattern worth knowing. They preserve the gist and drop the specifics, so anything that depends on exact values degrades first:
Numbers, dates, identifiers and prices, which summarisation quietly rounds away or omits.
Negations. "The customer did not approve the refund" compresses badly and reverses meaning when it goes wrong.
Anything the model must quote verbatim, such as legal or policy text.
Long-horizon conversations, where each round of summarisation compounds the loss from the previous one.
The mitigation is boring and effective: keep a structured, uncompressed slot for facts that must stay exact, and only compress the prose around it. A summary plus a small table of the entities involved beats a summary alone.
A sensible order to apply it
Delete dead prompt text. Free, no accuracy cost.
Turn on prompt caching if you have a stable prefix. Also no accuracy cost.
Retrieve rather than stuff, if you are sending whole documents.
Summarise old conversation turns, keeping exact values in a structured field.
Only then consider token-level compression, and measure it against your eval set rather than trusting a reported ratio.
Frequently asked questions
How much can you compress a prompt safely?
The pruning step is usually 20 to 40 percent with no measurable loss. Past that, every method trades accuracy, and the only honest answer is what your own eval set says.
Is prompt compression the same as summarisation?
Summarisation is one method of prompt compression. Pruning, retrieval and token-level compression are others, and they have quite different risk profiles.
Does compression reduce output cost too?
No. It reduces input tokens. Output length is governed by your instructions and the task. Since output tokens generally cost more, a short prompt requesting a long answer is still an expensive call. See what a token actually is for why the two are priced differently.
Should I compress if my prompt fits in the context window?
Fitting is not the only reason. Cost and time to first token both scale with input length regardless of headroom. But if the prompt is short and cheap already, leave it alone. What a context window is and the wider picture of how models work give the background if the sizing is unclear.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


