Dashboard

What Goes in the Prompt vs What Goes in Data

Deciding what goes in the prompt and what belongs in retrieved data comes down to two questions: does it change, and how big is it. Here is the rule.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
1 October 20261 min read

Deciding what goes in the prompt and what goes in the data is the architectural question people answer by accident. Something needs to be in context, it gets pasted into the system prompt, and six months later the prompt is 4,000 tokens of pricing tiers that change every quarter.

Two questions settle it. Does this change? How big is it? Everything else follows.

The rule: what goes in the prompt and what goes in the data

Small

Large

Stable

Prompt. Your tone rules, output format, worked examples

Prompt, if it fits and caches. Otherwise retrieval

Changes often

Inject per request. The user's name, today's date, the account tier

Retrieval. Product catalogue, docs, tickets, anything you would call a database

Stable and small is the only quadrant that genuinely belongs in the prompt text. The other three are either injected as variables at request time or fetched from somewhere and placed in context. The distinction matters because it decides who has to deploy to change a fact: a developer, or whoever owns the spreadsheet.

What belongs in the prompt

  • Role and scope. What the assistant is for, and what it declines.

  • Output format and tone. How answers are shaped.

  • Worked examples. One to three, demonstrating the format.

  • Decision rules that rarely change. Escalate anything mentioning a chargeback.

  • Refusal behaviour. What to do when it does not know.

The test: if changing this fact would require you to think about prompt wording, it is prompt content. If changing it is just updating a value, it is data.

What belongs in the data

  • Anything a non-engineer should be able to change without a deploy.

  • Anything with more than a handful of entries: products, policies, articles, past tickets.

  • Anything that would be stale within a week.

  • Anything per-user. Their order history belongs in the request, not in a prompt shared by every user.

The last one has a security dimension as well as an architectural one. A fact placed in a shared system prompt is visible to every conversation, and the model has no concept of which caller was supposed to see it.

Why a bloated prompt costs you twice

Putting data in the prompt fails in two ways at once, and the second is the expensive one.

First, you pay for it on every single call. A 4,000 token prompt sent 50,000 times a month is 200 million input tokens, whether or not any given request needed the pricing table.

Second, and worse, long stable prompts dilute attention. Instructions that matter compete with reference material that usually does not, and models follow instructions less reliably as the surrounding text grows. This is the mechanism behind a familiar complaint, covered in why your AI chatbot forgets earlier instructions.

A short prompt plus three retrieved paragraphs usually outperforms a long prompt containing everything, at a fraction of the token cost.

The middle ground: injected variables

Small and volatile is the quadrant people get wrong in the other direction, reaching for retrieval when a template variable would do:

python
SYSTEM = """You are the support assistant for Northwind Tools.
Answer only from the reference material provided. If it is not there, say
you will pass the question to a human and stop.

Today is {today}. The customer is on the {tier} plan.
"""

messages = [
    {"role": "system", "content": SYSTEM.format(today=today, tier=account.tier)},
    {"role": "user",   "content": f"<reference>\n{retrieved}\n</reference>\n\n{question}"},
]

Three things are happening there, and keeping them separate is the point. The stable instructions are prompt text. The two volatile facts are variables. The large volatile material is retrieved and clearly delimited. Anyone can change the retrieved corpus without touching code, and nobody has to redeploy because the date rolled over.

Where retrieval becomes the answer

The threshold is not a token count, it is a management question: can a person keep this correct by editing a prompt? Once the material has enough entries that nobody can hold it in their head, the prompt is the wrong home for it, even if it technically fits in context.

That is the practical argument for retrieval, and it arrives earlier than the context limit does. What RAG is in AI covers the mechanism, what chunking in RAG means covers the part that determines whether it works, and RAG vs fine-tuning vs long context compares the three ways of getting knowledge into a model.

A migration path for a prompt that already bloated

  1. Mark every line as stable or volatile. Read your prompt and label each fact. The volatile lines are your extraction list.

  2. Pull the large volatile blocks out first. Catalogues, policy text, pricing. Biggest token saving, lowest risk.

  3. Replace small volatile facts with variables. Dates, names, tiers, plan limits.

  4. Re-test before you delete anything. Keep the old prompt, run both against the same twenty inputs, compare. Shorter and worse is not a win.

  5. Leave the examples alone. Worked examples are stable and small, which is exactly where they belong.

Step 4 is not optional. Prompts accumulate load-bearing accidents, and some of that bloat is holding behaviour you have forgotten you rely on. The method is in how to test whether a prompt change actually improved your output.

Questions

Does prompt caching make a long prompt free?

Cheaper, not free, and only for the stable prefix. Cached input is typically discounted heavily, which is a strong argument for putting the stable parts first and the volatile parts last. It does not address the attention-dilution problem at all.

Should retrieved data go in the system or user message?

The user message, delimited with tags, below the question or above it consistently. It is material for this request, not a standing instruction, and keeping it out of the system prompt preserves the cacheable prefix.

How do I stop the model treating retrieved text as instructions?

Delimit it clearly and say in the system prompt that reference material is information, never instructions. Treat that as partial mitigation rather than a fix: a determined injection in retrieved text can still win, which is why the delimiting is a habit and not a defence.

Is there a size where a prompt is definitely too long?

No universal number, but if your system prompt is longer than the typical user question by more than a factor of ten, look at it. The useful signal is content, not length: count how many lines would need editing if a price changed.

Where should I start if all of this is new?

Prompt engineering is the overview, and it covers the container most of this lives in before it reaches the architecture question above.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.