Dashboard

How to Prompt AI to Write Unit Tests That Catch Bugs

AI-generated unit tests often pass without testing anything real. The fix is stating the requirement before the AI ever sees the code.

Steve Jefferson
Steve Jefferson
Developer Advocate
15 September 20261 min read

How to Prompt AI to Write Unit Tests That Catch Bugs

The fastest way to get unit tests that catch real bugs is to tell the AI what the code should do before it sees how the code currently does it, then ask it to write tests against that behavior spec, not against the implementation. Skip that step and you get the single most common failure in AI-generated tests: assertions that check whatever the code currently outputs, which pass immediately and catch nothing, ever, because they were never testing a requirement in the first place.

The failure mode, made concrete

Ask an AI coding agent to "write tests for this function" with no other context, and it will read the function, run it in its head, and write assertions that match what it saw. If the function has a bug, the test enshrines the bug as correct behavior. The test suite goes green. Nothing is protected.

Here is what that looks like in practice. A function meant to calculate a discount incorrectly applies it before tax instead of after. Asked to "write tests for calculateDiscount," a model with no behavior spec will write:

text
test('calculates discount correctly', () => {
  expect(calculateDiscount(100, 0.1, 0.08)).toBe(97.2);
});

That number is whatever the buggy function actually returns. The test passes today, and it will keep passing after someone else "fixes" the bug and the number changes to something the test now calls wrong. The test was never protecting a requirement. It was recording an accident.

The prompt pattern that fixes it

Separate the spec from the implementation, explicitly, every time:

  1. **State the requirement first, in plain language, before showing the code.** "Discount should apply to the pre-tax subtotal, then tax should be calculated on the discounted amount." Now the model has something to check the code against, not just something to describe.

  2. **Ask for the test to be written from the requirement, then run against the code.** "Write a test that verifies this rule. If the current implementation fails it, tell me, do not silently write a passing test."

  3. **Ask explicitly for edge cases the happy path won't surface**, since a model prompted only for "tests" tends to write one obvious case and stop: zero values, negative inputs, boundary values (a 100% discount, a $0 order), and null or missing fields.

  4. **Ask for at least one test that should fail given the current code**, if you suspect a bug. A model that only ever writes passing tests is optimizing for green, not for correctness.

A prompt template you can reuse

text
Requirement: [plain-language spec of correct behavior, written by a human,
not derived from reading the code]

Write unit tests for [function/module] that verify this requirement.
Do not infer the expected behavior from the current implementation.
If the current code fails any test you write, say so explicitly rather
than adjusting the test to match.

Include:
- the documented happy path
- at least 2 boundary/edge cases
- at least 1 case that exercises invalid or missing input

This works because it removes the model's only path to a tautological test: reading the code to figure out what "correct" means. Once the requirement is stated first, the code becomes something to check, not something to copy.

A full before-and-after example

Take a real function: a shipping cost calculator that is supposed to give free shipping on orders over $50, otherwise charge a flat $5.99.

Prompted with just "write tests for this function," a model that reads a buggy version (one that compares against 50 instead of >= 50, so a $50.00 order gets charged) produces:

text
test('charges $5.99 under the free shipping threshold', () => {
  expect(shippingCost(49.99)).toBe(5.99);
});
test('charges $5.99 at exactly $50', () => {
  expect(shippingCost(50.00)).toBe(5.99);
});

Both pass. Both are wrong to have, because the second one encodes the bug as the spec.

Prompted with the requirement stated first ("free shipping applies at $50 and above, inclusive"), the same model produces:

text
test('charges $5.99 under the free shipping threshold', () => {
  expect(shippingCost(49.99)).toBe(5.99);
});
test('gives free shipping at exactly $50, inclusive', () => {
  expect(shippingCost(50.00)).toBe(0);
});
test('gives free shipping above $50', () => {
  expect(shippingCost(75.00)).toBe(0);
});

The second test now fails against the buggy implementation, exactly as it should. That single failing assertion is the entire value of writing the requirement down first: it turns an invisible off-by-one bug into a red test, the moment the tests are written rather than whenever a customer notices they got charged for a $50.00 order.

A manual sanity check that catches what the prompt misses

Even a good prompt occasionally produces a test that would pass no matter what the code does, usually because it asserts on a type or existence rather than a value ("expect result to be defined"). The fastest manual check: temporarily break the function on purpose, something obviously wrong, and re-run the test suite. Any test that still passes is not testing anything. This is a simplified version of mutation testing, and running it once by hand on a batch of AI-generated tests takes minutes and reliably surfaces the fakes.

Where this connects to the rest of your workflow

Writing tests from a stated behavior spec pairs naturally with getting an AI coding agent to reproduce a bug before fixing it: both practices force the model to establish what correct or broken actually means before it starts changing code, rather than pattern-matching its way to something that looks plausible. If you are deciding between letting a model regenerate a function outright versus patch it, fix or regenerate covers that decision directly, and the same requirement-first discipline applies to reviewing whichever path you pick.

For teams building out a broader prompting practice, this is one instance of a general rule: a complete list, not a sample covers the sibling failure mode, where a model gives you a representative subset when you needed exhaustive coverage. Tests and lists fail the same way, for the same underlying reason: the model defaults to "plausible" unless you specify "complete" or "correct against a stated rule."

Frequently asked questions

Can AI write good unit tests without any human input?

Not reliably. A model with no stated requirement will write tests that match the current implementation, bugs included. The requirement is the one piece of context a model cannot infer correctly on its own, since it depends on intent, not on the code in front of it.

How many edge cases should I ask for?

Enough to cover zero, negative, boundary, and missing-input cases at minimum for numeric or user-facing functions. Ask for more once you know your domain's specific failure history, past bugs are the best source of edge cases to test for going forward.

Should I ask the AI to write tests before or after it writes the implementation?

Before, when practical. Writing the test from the requirement first, then generating the implementation against that test, is the AI-assisted version of test-driven development, and it structurally prevents the tautological-test problem since the test exists before there is any implementation to copy from.

What if I don't know the exact requirement, I just want the code tested?

Write down what you do know, even loosely, "should never return a negative price" is a real requirement. A vague or partial spec still gives the model something to check against, which is strictly better than nothing. If you truly cannot state any requirement, that is a sign the code's purpose needs clarifying before it needs tests.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.