Dashboard

How to Prompt AI for a Complete List, Not a Sample

Ask AI to list everything and it quietly stops at eight of nineteen. Count first, chunk the input, and verify adversarially to get the full set.

Steve Jefferson
Steve Jefferson
Developer Advocate
12 September 20261 min read

How to Prompt AI for a Complete List, Not a Sample

Ask a model to list every deprecated function in a file and it will hand you eight, confidently, when there are nineteen. Nothing in the answer says it stopped early. There is no ellipsis, no "and others", no hedge. It looks finished because a good answer to most questions is a representative one, and completeness is a different task wearing the same sentence.

The fix is not a firmer instruction. It is to stop asking for a list and start asking for a count, a pass over bounded chunks, and a verification step.

Why "list all" produces a sample

Three things work against you at once.

  • A model generates the most plausible continuation, and after eight good items a ninth becomes less likely than a closing sentence. The pull toward a well-formed ending is strong, and it does not care whether you finished.

  • There is no internal counter. Nothing tracks how many candidates existed or how many have been emitted, so nothing notices the gap.

  • Long outputs are the expensive, slow half of a request. Everything about how models are built and tuned nudges toward concision, which is usually right and is exactly wrong here.

Notice that none of these are fixed by adding "ALL" in capitals, promising a tip, or saying it is very important. The failure is structural. Instructions do not change structure.

Count first, then enumerate

The single highest-value change is to make the model commit to a number before it starts writing items. A count is cheap to produce, and once it exists, a list of eight against a stated nineteen is a visible contradiction rather than a silent omission.

text
Step 1. Count only. Read the file and reply with a single number:
how many function definitions use a deprecated API. No list, no
explanation, just the number.

Step 2. (new message) There are 19. List all 19, numbered 1 to 19,
one line each, in the order they appear in the file. If you reach the
end of the file before 19, say so and stop.

Two things make this work. The number arrives before the generation pressure builds, and the numbered format means the model is producing "14." before it produces the item, which is a much stronger commitment to continue than a bullet point.

The explicit permission to say it fell short matters as well. Without it, a model that cannot find nineteen will invent the difference. How to make AI say I don't know covers why that escape hatch has to be stated.

Chunk the input so completeness is possible

If the source is long, no amount of prompting will get you a complete list in one pass, because attention over a large document genuinely degrades toward the middle. Split it.

  1. Divide the source into pieces small enough that each is comfortably readable. A few thousand tokens each is a reasonable working size.

  2. Run the same extraction prompt on each piece independently, in its own request. Independence is the point: one long conversation reintroduces the same pressure you just removed.

  3. Ask for a count per piece as well as the items.

  4. Concatenate, then deduplicate across boundaries, where the same item can legitimately appear twice.

  5. Sum the counts and compare with the length of your merged list. A mismatch is a specific piece to re-run, not a vague doubt about the whole job.

This is slower and more expensive than one clever prompt, and it is the difference between a list you can act on and a list you have to check by hand anyway. How to chain prompts together covers the plumbing.

Verify by asking the opposite question

A second pass framed as adversarial catches more than a second pass framed as review. Asking "did you miss anything" invites a polite no. Asking the model to find a specific fault invites it to look.

Useful framings, in rough order of effectiveness:

Framing

What it catches

Here is my list of 19 and the source. Find items in the source that are not on the list.

Straight omissions, the main failure

For each item on the list, quote the line it came from.

Invented items, which appear when the count was too high

A reviewer says this list is missing at least three items. Find them.

Omissions the model was reluctant to admit

Which category of matching item is this list weakest on?

Systematic blind spots rather than one-off misses

Run the verification in a fresh context, without the conversation that produced the list. A model shown its own prior work tends to defend it. How to prompt AI to check its own work goes further on that effect.

When to stop prompting and write code

Worth saying plainly: if the thing you want is mechanically findable, find it mechanically. Every deprecated function call, every TODO, every email address in a directory. A regular expression is complete by construction, runs in a second, and does not need verifying.

Use the model for the judgement layer that has no regular expression: which of these matter, what category each belongs to, which are false positives. Extraction by code, classification by model, is both cheaper and more reliable than asking one system to do both. The rest of our prompt engineering coverage works the same seam.

FAQ

Does a bigger context window fix this?

It removes the excuse, not the behaviour. A model that can see the whole document still prefers a tidy answer to an exhaustive one, and recall across a very long input is uneven. Does a bigger context window mean better answers covers the wider point.

Should I raise max_tokens?

Check it is not the constraint, then stop thinking about it. Truncation at the token limit looks different, usually cutting off mid-item. A clean, well-formed short list is the model choosing to finish, not the API stopping it.

Does asking for JSON help?

Somewhat. A structured format with an explicit total field alongside the array gives you the count and the items in one response, and the mismatch becomes machine-checkable. It does not fix recall, it just makes the shortfall detectable without you reading anything. How to get JSON output from AI covers reliable structure, and OpenAI's structured outputs documentation covers schema enforcement where the provider supports it.

What if I do not know the true count either?

Then the count step is still useful as an internal consistency check, and the adversarial pass becomes the load-bearing one. Run the extraction twice in separate contexts and compare. Items appearing in one run and not the other tell you where the uncertainty is concentrated.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.