Prompt AI With a Worked Example, Not Rules
Prompt AI with a worked example and it beats a list of rules. How to build one, what a good example must contain, and the cases where examples hurt.
When you prompt AI with a worked example instead of a list of rules, you stop describing the output and start showing it. One complete example usually beats six paragraphs of instructions, because an example encodes the things you would never think to write down. Length control works the same way, see how to prompt AI to give a shorter answer.
Here is the same task both ways, so the difference is visible rather than asserted.
The instruction version
Summarise this customer support email. Be concise. Use a neutral tone.
Include the customer's main issue and what they want. Do not include
pleasantries. Keep it under 40 words. Use sentence case. If they mention
an order number, include it. Do not speculate about causes.Eight rules. The output will respect most of them and will still surprise you: it will open with The customer, or write the order number as Order #4471 when your system expects 4471, or run to 38 words of which 12 are scaffolding.
Prompt AI with a worked example instead
Summarise support emails like this.
EMAIL:
Hi there, hope you're well! I ordered a desk lamp two weeks ago (order
4471) and it still hasn't turned up. Tracking says delivered but nothing
arrived. I'd really rather just have a refund at this point. Thanks, Dani
SUMMARY:
Order 4471 marked delivered but not received. Requests a refund.
Now do this one:
EMAIL:
...The example carries every rule the instruction version listed, plus three it did not: that the order number is bare, that the summary leads with the fact rather than the customer, and that a request is phrased as a verb rather than a sentence about wanting something. Nobody would think to write those down. The example states them without saying them.
What a useful example has to contain
A bad example is worse than none, because the model will copy whatever is actually in it, including your accidents. Four properties:
It must be a real case, not an idealised one. Pick a genuine input from your data, with its mess intact. An example built from a clean imaginary input teaches the model to expect clean input.
The input and output must be visually separated. Use consistent labels like EMAIL: and SUMMARY:, or XML-style tags. The model is learning the boundary as well as the content.
The output must be exactly what you want, down to the punctuation. If your example ends without a full stop, expect outputs without full stops. This is a feature, and it is how you control format without describing it.
It must be representative, not exceptional. An example of your weirdest edge case teaches the model that edge cases are the norm.
How many examples
More is not better past a small number, and the curve is steeper than people expect:
Examples | What it gets you | When to stop here |
|---|---|---|
0 | Baseline. The model's default interpretation | Simple tasks where the default is already right |
1 | Format, tone and structure locked in. The largest single jump | Most formatting and transformation tasks |
2 to 3 | Teaches variation: how to handle a case with no order number, or a complaint rather than a request | Tasks where inputs differ in kind, not just content |
4 or more | Diminishing returns, rising token cost, and a growing risk of the model over-fitting to surface patterns | Rarely worth it; prefer fixing the examples you have |
If three examples are not enough, the problem is usually that the task contains two tasks. Split it. The broader comparison is covered in few-shot vs zero-shot prompting.
Use your examples to define the edges
The highest-value second example is not another typical case. It is the case where the rule changes. If an email has no order number, what then? Show it:
EMAIL:
Your checkout page won't accept my card, tried three times. Chrome on a Mac.
SUMMARY:
Checkout rejects card payment. Chrome on macOS, three attempts.That one example resolves a question no amount of instruction phrasing settles cleanly: the summary simply omits the order number rather than writing none or N/A. You have defined behaviour at the boundary by demonstration.
The same trick works for refusals. If some inputs should produce no output, show an example where the output is a specific marker, and the model will use that marker rather than inventing an apology.
When examples make things worse
Three cases where reaching for an example is the wrong move:
Open-ended generation. Asking for ten product name ideas with one example anchors every suggestion to that example's shape. Here, describing constraints beats demonstrating one answer.
Reasoning tasks where the path matters more than the format. An example of a finished answer can teach the model to jump to a conclusion in the same shape without doing the work. Ask for the reasoning instead, as in how to prompt AI to explain its reasoning before it answers.
When your examples disagree with each other. Two examples that handle the same situation differently are worse than one, because you have demonstrated that the rule is arbitrary.
The second case is the subtle one. A worked example is excellent at teaching form and mediocre at teaching thought, and a lot of disappointing results come from using it for the second job.
The counterpart: showing what you do not want
Everything above is about demonstrating the target. The mirror technique, showing an example of the output you are trying to avoid, works too and has its own rules, because a naive bad example anchors the model to the thing you just showed it. Prompting AI with an example of what you do not want covers the contrastive pairing that makes it work.
The short version: a negative example is only safe next to a positive one, labelled, in that order. On its own it reads as a demonstration.
Where to put the example
If you are using a system prompt, examples belong there, not in the user turn. They are instructions about how to behave, they do not change between requests, and keeping them out of the user message means they survive a long conversation rather than scrolling out of attention. System prompt vs user prompt covers the division, and how to write a system prompt for a custom AI assistant covers the surrounding structure.
Stable examples in a system prompt are also the part of your prompt most likely to be cacheable, which matters once volume is real.
Writing the example from an output you already liked
The fastest way to build a good example is not to write one. It is to go back through your own history, find a response you were happy with, and pair it with the input that produced it.
That pairing is already calibrated to your taste in a way a freshly written example is not, and it takes two minutes. Three things to do to it before you keep it:
Trim it to the shortest version that still demonstrates the format. The response you liked is probably longer than it needs to be, and every extra token ships on every call.
Strip anything specific to that one request. A reference to a particular customer or date teaches the model that those details belong in the output.
Check the input is typical. You remember the output. Look at what went in and ask whether it resembles your normal traffic.
Keep the originals somewhere, because this is how a prompt library accumulates: a handful of input and output pairs you trust is a more durable asset than the prompts built on top of them.
Check that the example is doing the work
Before you keep an example, prove it earns its tokens. Run twenty real inputs with it and twenty without, grade both against what you actually wanted, and compare. If the scores match, delete the example and keep the shorter prompt.
This is the same discipline as any other prompt change, and the method is in how to test whether a prompt change actually improved your output.
Questions
Should the example use real customer data?
Use a real case with the identifying details changed. Keep the structure and the mess, replace names, emails and numbers. A prompt is logged by your provider and read by your team, so it is not the place for personal data.
Do examples count against my token budget on every call?
Yes, every single call. That is the cost side of the trade, and it is why two good examples beat five mediocre ones.
My model ignores the example format. Why?
Usually because the example and the instructions contradict each other. If the instruction says under 40 words and the example is 50, the model has to pick. Delete the instruction and let the example rule.
Does this work for structured output like JSON?
It works, and a schema works better. Use the provider's structured output mode where it exists, which guarantees shape, and keep the example for the semantic choices a schema cannot express, such as how terse a field should be.
Is this the same as fine-tuning?
No. Examples in a prompt are read at request time and cost tokens; fine-tuning changes the model's weights and costs a training run. Start with examples, because they are reversible in a second.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


