Turn a Long Email Thread Into a Decision With AI
Turn an email thread into a decision with AI. Summarising gives a recap, not an answer. A prompt that extracts decisions and open questions, with quotes.
Forty-one messages, nine people, three weeks, and somewhere in there the question of whether you are shipping in November. Pasting the whole thread into a chat window and asking for a summary gives you a summary, which is the one thing you do not need. You already know roughly what it says. What you need is the decision, who owns it, and what is still open.
To turn an email thread into a decision with AI you have to stop asking for a summary. The difference is almost entirely in the prompt, and here is how to write one that produces a decision rather than a recap.
Why summarise fails here
A summary prompt optimises for coverage. It tries to represent the whole thread proportionally, which means the twelve messages arguing about a date get twelve messages' worth of weight, and the one message where someone quietly agreed to the date gets one.
Decisions are not proportional. They are usually one or two sentences buried in the middle of a long message, often phrased tentatively, frequently never restated. "I suppose we could live with the 14th" is the decision. It will not survive summarisation, because summarisation is averaging and that sentence is an outlier.
There is a second problem specific to long threads. Models reliably attend better to the start and the end of a long input than the middle, the effect known as lost in the middle. An email thread puts the oldest messages first, the newest last, and the moment where things actually got decided somewhere in between. The structural weakness of the model lines up exactly with where the answer lives.
Both problems have the same fix: stop asking for a representation of the thread and start asking for specific extractions with evidence.
The prompt that turns an email thread into a decision
Below is an email thread. Do not summarise it.
Produce exactly these four sections:
DECISIONS MADE
For each decision: what was decided, who decided it, the date,
and a direct quote of the sentence where it was decided.
If a decision was later reversed, show both and mark it reversed.
STILL OPEN
Questions raised that nobody answered. For each: who raised it,
what they asked, and whether anyone acknowledged it.
COMMITMENTS
Anything anyone said they would do. Name, action, stated deadline
if any. Include vague ones and mark them vague.
DISAGREEMENTS
Points where two people stated incompatible positions that were
never resolved. Quote both sides.
Rules:
- Every item must have a quote. No quote, no item.
- Do not infer agreement from silence.
- If a section is empty, write "None found" rather than filling it.
- Flag anything where you are unsure whether it counts.
Thread:
[paste]Four things in that prompt do the work.
The quote requirement is the main one. It forces the model to point at the actual sentence, which both grounds the output and makes it checkable in seconds. You read the quote, not the characterisation.
"Do not infer agreement from silence" blocks the most common and most expensive error. Threads are full of proposals nobody responded to, and a model asked to find decisions will happily promote a proposal to a decision because it was not contradicted. In an email thread, silence means people were busy.
"None found" rather than filling it stops the model manufacturing content for an empty section. Without this instruction, asking for four sections reliably produces four populated sections.
The vague and unsure flags give it somewhere to put the genuinely ambiguous material, which is otherwise forced into a definite category. Prompting a model to flag what it is unsure about works much better when there is an explicit place for the uncertainty to go.
Dealing with the middle of long threads
For anything over about twenty messages, do not paste it as one block.
Split it into chunks of five to eight messages and run the same prompt on each, then run a final pass over the combined output. This costs more calls and takes longer, and it is noticeably more accurate, because every message gets to be near the edge of some context window rather than buried in the middle of one.
Keep the dates and senders on every chunk. A decision without a date cannot be ordered against the decision that reversed it, and reversal is the thing you most need to get right.
Reading the output properly
Check the quotes first. Pick two or three items and search for the quoted sentence in the original thread. If a quote is not there verbatim, or is there but says something different in context, discard the whole output and rerun. This takes ninety seconds and catches the failure mode that matters.
Then read STILL OPEN before DECISIONS MADE. The open questions are the actionable part. The decisions you probably half-remember; the question somebody asked on day four that nobody answered is the one that will surface as a problem later.
Treat DISAGREEMENTS with suspicion in both directions. Models under-report disagreement because polite professional language disguises it, and occasionally over-report it by reading a clarifying question as opposition. Verify both ways.
Making it stick
The output is for you, not for the thread. Resist pasting it back in: a summary of a thread posted into the thread is how threads reach fifty messages.
Instead, send one short message with the decisions you believe were made, the open questions, and a request to correct anything wrong. That is a different act. It forces the implicit into the explicit and gives people something cheap to respond to. If a decision was never really made, this is when you find out, which is the entire point of the exercise.
If the thread involves a client and the next step is a follow-up rather than a decision log, prompting AI for a follow-up email after a sales call covers that adjacent case.
When to use a different approach entirely
If the thread is short, under about ten messages, read it. The setup cost of doing this well exceeds the reading time, and you will have better judgement about tone than any extraction will.
If the thread is contentious, read it yourself and use the extraction as a check on your reading rather than a replacement. Conflict is exactly where quoting out of context does the most damage, and where you most need to have formed your own view.
If you suspect the model is telling you what you want to hear about where things landed, prompting it to disagree with you is a useful second pass: state your reading of the outcome and ask it to argue the opposite from the same thread.
For the general techniques underneath all of this, the prompt engineering guide is the broader reference.
FAQ
Should I remove the email signatures and quoted replies first?
Yes, where it is easy. Quoted reply chains duplicate content and make the model see the same sentence many times, which inflates its apparent importance. Signatures are harmless noise but add tokens.
How do I handle a thread that forked into several threads?
Run each fork separately, then run the final combining pass over all of them together. Forks usually diverge on exactly the point that was never resolved, so the comparison is informative.
Can I trust it with a thread containing confidential information?
That depends on your tool and your data agreement, not on this technique. Check where the text goes before pasting a client thread anywhere.
What if the decision was made in a meeting and only referenced in the thread?
The extraction will surface the reference, which is the useful outcome. "Per Tuesday's call we're going with option B" shows up as a decision with a quote, and tells you the authority sits in a call you need the notes from.
Why not just ask for action items?
Action items collapse four different things into one list. Commitments, open questions, decisions and disagreements need different responses from you, and merging them means the open questions quietly become somebody's to-do rather than something anyone answers.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


