How to Prompt AI to Write Test Cases That Matter
The naive test-writing prompt produces tests that only confirm code does what it obviously does. Here is the two-pass structure that fixes that.
Ask a model to "write test cases for this function" and you get a pile of tests that all check the function does what it obviously does, plus maybe one null check. That is not what test cases are for. Good tests exist to catch the ways code breaks that you did not think of while writing it, and a model producing tests from the same mental model as the code that inspired the request will miss exactly the same blind spots the developer did.
Why the Naive Prompt Fails
The core problem is that "write tests for this" gives the model no information about what could go wrong, only about what the code is supposed to do. A model reading a function that validates an email address will happily generate three tests confirming that valid-looking emails pass and invalid-looking ones fail, and stop there, because nothing in the prompt told it to think adversarially. The fix is not a better one-line prompt, it is a structured request that makes the model enumerate failure categories before it writes a single test.
A Two-Pass Prompt Structure
Split the request into two explicit passes: first enumerate what could go wrong, then write tests against that list. Doing both in one shot tends to collapse back into the naive pattern, because the model writes tests as it thinks of failure modes rather than against a complete list.
Pass 1: Before writing any tests, list the categories of failure
this function or feature could have. Include at minimum: invalid
or malformed input, boundary values, empty or null cases, concurrent
or repeated calls, failure of any external dependency it relies on,
and at least one case where a caller uses the function in a way its
name suggests is fine but its actual contract does not support.
Do not write test code yet, just the list.
[paste function or feature description]Review that list before moving to pass two. This is the moment a human catches the categories that do not apply (no external dependency here) and the ones the model missed (nothing about internationalization, if the function handles user-facing text). Then:
Pass 2: Write test cases for each item in the list above. For
each test, name it after the specific behavior being verified,
not a generic "test case 1" label, and include the expected
outcome in a comment above the assertion. Skip categories from
the list that do not apply, but say explicitly which ones you
skipped and why.The instruction to name tests after specific behavior matters for a reason beyond readability: a model asked to produce well-named tests tends to actually think through what each test proves, rather than generating structurally similar tests that differ only in input values.
Feeding It the Contract, Not Just the Code
Tests are only as good as the model's understanding of what the code is supposed to do, and code alone under-specifies that. If a function has a docstring, ticket description, or even a Slack message explaining the intended behavior, include it. Without that context, a model can only test what the code currently does, which means a real bug (code that runs without error but does the wrong thing) gets a passing test written around it instead of a failing one that catches it.
This is the single biggest difference between test cases that just increase a coverage percentage and test cases that catch real bugs: coverage percentage only requires exercising the code path, while catching a genuine bug requires the model to know what correct behavior looks like independently of what the code currently does.
Testing the Failure Paths Specifically
For anything that calls an external service, whether that is a payment API, a database, or another internal service, explicitly ask for tests that simulate that dependency failing, timing out, and returning malformed data, not just succeeding. This category is the one developers skip most often when writing tests by hand, because it requires imagining infrastructure failures rather than just business logic, and it is exactly the category where a structured prompt earns its keep.
Failure path | Common gap in hand-written tests |
|---|---|
Dependency times out | Usually untested; code often has no explicit timeout handling to verify |
Dependency returns malformed data | Tested for the happy shape, not for a truncated or wrong-typed response |
Dependency is unavailable | Tested as "throws an error," not as the specific retry or fallback behavior expected |
Two calls race | Rarely tested at all outside of dedicated concurrency-critical code |
Reviewing What the Model Produces
Treat AI-generated tests the same way you would treat a junior engineer's first pass: read every assertion, not just the test names. A model can write a test that looks thorough but asserts the wrong thing, most commonly asserting that code returns without throwing rather than asserting the actual return value is correct, which passes even if the underlying logic is subtly broken. Run the generated tests against a deliberately broken version of the code (comment out a validation check, flip a boundary condition) before trusting them; a test suite that still passes against broken code is not testing anything.
Watch for Tests That Are Flaky by Design
A model asked to test anything involving time, randomness, or concurrency will sometimes write an assertion that is technically correct but non-deterministic in practice, a test that checks a timestamp is "recent" without a tolerance window, or one that asserts an exact ordering from a process that does not guarantee one. These tests pass in development and fail intermittently in CI, and because they usually fail on the tests they generated, not the ones a human reviewed carefully, they are one of the more expensive traps in AI-generated test suites: a flaky test erodes trust in the whole suite faster than a missing test does, because people start ignoring red CI runs.
When reviewing generated tests, specifically scan for anything asserting on wall-clock time, random values, or network timing without an explicit tolerance or mock, and either fix it or send it back with a pointed instruction: "this assertion depends on timing; rewrite it to be deterministic, using a fixed clock or a tolerance range." Models correct this readily once told, but rarely avoid it unprompted, since a plausible-looking assertion about "roughly now" reads as correct on its own without the context of running in a shared CI environment under load.
Frequently Asked Questions
Should I use this same approach for integration tests, not just unit tests?
Yes, with one addition: for integration tests, explicitly list the real systems involved and ask the model to identify which failure categories are relevant to each boundary between systems, since integration bugs usually live at those boundaries rather than inside any single function.
How many tests is too many for one function?
There is no fixed number; the right stopping point is when the failure-category list from pass one is exhausted, not when a round number of tests has been reached. A short, well-targeted list of five tests covering five real failure modes beats twenty tests that are minor variations on the same happy path.
Does this work for testing AI features themselves, like a prompt or an agent?
The failure-category framing still applies, but the categories differ: think ambiguous instructions, adversarial or off-topic input, and cases where the correct behavior is refusing to answer, rather than boundary values and null checks in the traditional sense.
Related Reading
Once you have a strong test suite, the natural next problem is what happens when an AI coding agent needs to fix a memory leak or other bug the tests surface. For choosing which model to lean on for this kind of structured technical prompting in the first place, see how to benchmark an AI coding agent on your own codebase. Auditing your dependencies with AI covers a related structured-prompting pattern for a different code-quality task, and the broader technique library lives in prompt engineering: a framework for prompts that actually work.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


