What Is Prompt Chaining (And When to Use It)
A single messy task, run as one prompt versus a three-step chain, shows exactly where the one-shot version quietly loses items the chain catches every time.
Prompt chaining means breaking a task into a sequence of smaller prompts, where each step's output becomes the next step's input, instead of asking a model to do the whole thing in one shot. It sounds like more work. For anything with more than one real step buried inside it, it is usually less work, because it catches mistakes a single prompt lets through silently.
The test case: a messy meeting transcript
Take a rough, unedited transcript of a 40-minute planning meeting and ask for a structured list of action items with owners and deadlines.
One big prompt
"Read this transcript and extract all action items with who owns each one and when it is due." This works, mostly. It also tends to miss items that were mentioned in passing rather than stated as a clear commitment, merge two separate action items that got discussed back to back, and guess at an owner when the transcript only implied one through context rather than stating a name directly. None of these failures throw an error. The output looks complete. It is just quietly short.
A three-step chain
Step one: "List every sentence in this transcript that could plausibly be a commitment, a task, or a next step, even a vague one. Do not filter yet, over-include."
Step two, fed the output of step one: "For each item in this list, state the specific action, the owner if named or clearly implied, and the deadline if mentioned. Mark owner or deadline as unclear rather than guessing if it is not stated."
Step three, fed the output of step two: "Merge any duplicate or overlapping items from this list into a single clean action item each, and flag any item still missing an owner or deadline for human review."
The chain costs three model calls instead of one. What it buys back: step one's deliberate over-inclusion means nothing gets silently dropped before it is even considered, step two forces an explicit unclear marker instead of a confident guess, and step three catches the duplicates that a single pass tends to merge incorrectly or miss entirely. Run both versions against the same transcript and the chain's output reliably has two to four more real action items than the one-shot version, mostly the ones that were implied rather than stated outright.
Why this works, mechanically
A single large prompt asks a model to do extraction, judgment, and cleanup simultaneously, and those three jobs compete for the same pass of attention over the same input. Splitting them into separate steps means each step only has to do one job well, and the next step gets to work from a cleaner, already-processed version of the problem instead of the raw, messy original. This is the same reason few-shot versus zero-shot prompting matters: constraining what a single call is responsible for tends to improve the reliability of that call, whether the constraint comes from examples or from scope.
When chaining is not worth it
Not every task benefits. If the task is genuinely one step (summarize this paragraph, translate this sentence), chaining adds latency and cost for no reliability gain. The tell that a task actually needs chaining: can you name two or more distinct jobs happening inside the single request you were about to write? If yes, those are probably your chain's steps. If you cannot name a second job, it is not a chaining candidate, it is just one prompt.
Where the instructions for each step should live
Each step in a chain can carry its own system prompt scoped to that step's single job, which is a natural fit with how system prompts differ from user prompts: the system prompt sets the step's narrow role, the user prompt carries that step's actual input. For the fundamentals of writing any of these prompts well before you start chaining them, see our prompt engineering framework.
FAQ
Does prompt chaining cost more?
Yes, proportionally to the number of steps, since each step is a separate model call. For tasks where reliability matters more than the marginal cost of an extra call, it is usually worth it.
Is prompt chaining the same as an AI agent?
No. A chain is a fixed, predetermined sequence of steps you designed in advance. An agent decides its own next step based on what happened in the previous one, which is a more dynamic and less predictable process than a chain.
How many steps should a chain have?
As many as there are genuinely distinct jobs in the task, and no more. Three is common for tasks like the example above. Adding steps that do not correspond to a real distinct job just adds latency without improving results.
Can I chain across different models?
Yes, and it is a common pattern: use a cheaper, faster model for a high-volume early step like the over-inclusive extraction pass, and a stronger model for a later step that needs more judgment, like the merge and review pass.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


