Dashboard

How to Prompt AI to Write a Database Seed Script

A practical prompt template for getting an AI coding agent to write a seed script that respects foreign keys, generates realistic data, and survives a second run.

Steve Jefferson
Steve Jefferson
Developer Advocate
5 September 20261 min read

Getting an AI coding agent to write a database seed script that actually works takes more than asking it to "add some sample data." You need to hand it the schema, spell out the foreign key insert order, ask for skewed distributions instead of uniform ones, and require idempotency so the script survives a second run without errors. Below is a copy-paste prompt built around a real three-table schema, users, products, and orders, plus the three failure modes that show up in almost every AI-written seed script and the exact instruction that fixes each one.

What a seed script actually has to do

A seed script is not the same job as a schema migration. A migration changes the shape of the database: new tables, columns, constraints. A seed script fills an already-shaped database with rows so a developer, a QA tester, or a demo environment has something to look at. If you're still deciding how the tables should look, prompting AI to design the schema itself is a separate problem worth solving first, and if AI is also writing your database migrations, the failure modes there run more toward destructive changes and rollback safety.

Whichever AI coding tool you're using to write it, the seed script itself needs to do three things reliably: insert rows in an order that respects foreign keys, generate values that look like production data instead of test fixtures, and run safely more than once. Most AI-written seed scripts look right on the first read and fail on at least one of these the moment you actually run them.

The example schema this guide uses

Three tables cover almost every seeding problem you will hit in practice.

  • users: id (uuid), email (unique), full_name, created_at

  • products: id (uuid), name, price_cents, category

  • orders: id (uuid), user_id references users, product_id references products, quantity, status, created_at

orders depends on both of the other tables. That single dependency is enough to expose foreign key ordering bugs, and the three tables together are enough to show what a realistic distribution should look like: not every user has the same number of orders, and not every product costs a round number.

A copy-paste prompt for an AI-generated seed script

Paste your real schema, or the CREATE TABLE statements, in place of the example below, adjust the row counts, and hand the whole thing to your AI coding agent.

Write a Node.js seed script (using [your ORM or driver, e.g. Prisma with PostgreSQL]) that inserts sample data for this schema:

- users: id (uuid), email (unique), full_name, created_at
- products: id (uuid), name, price_cents, category
- orders: id (uuid), user_id -> users.id, product_id -> products.id, quantity, status, created_at

Requirements:
1. Insert in dependency order: users and products first, then orders. Use the actual ids you just inserted, never hardcoded or guessed ids.
2. Generate 200 users, 40 products, 600 orders. Use realistic, varied values, not sequential placeholders like "User 1", "User 2".
3. Make the distribution uneven on purpose: most users should have 0-2 orders, a small group (10-15%) should have 10 or more. Vary order status across pending, shipped, delivered, and cancelled, with delivered as the most common.
4. Make the script idempotent: running it twice must not create duplicate rows or throw a unique constraint error. Use upsert or ON CONFLICT DO NOTHING keyed on a natural unique field, or delete previously seeded rows before reinserting.
5. Wrap the inserts in a transaction so a failure partway through does not leave the database half-seeded.
6. Print a summary count of rows inserted per table when it finishes.

That prompt works whether the agent reaches for Faker.js, a similar data-generation library, or hand-written value lists, because the requirements describe the behavior you need, not the library you happen to name.

Three failure modes in AI-generated seed data, and the fix for each

Foreign key insert order breaks the script

The most common first failure: the agent writes the orders insert before the users and products inserts finish, or it hardcodes ids it assumes already exist. Either way, you get a foreign key violation on the first run.

The fix: state the dependency order explicitly and forbid hardcoded ids. "Insert in dependency order: users and products first, then orders. Use the ids you just inserted, never hardcoded ids" removes the ambiguity that causes this. If your schema has more than two or three levels of dependency, list the full order table by table rather than trusting the agent to infer it from the schema alone.

Unrealistic, uniform data distributions

Left alone, an AI coding agent tends to spread data evenly: every user gets exactly three orders, every price ends in .99, every signup date sits an identical distance from the next. That is close to useless for testing pagination or any query that depends on skew, like finding the top 10 percent of customers by order count.

The fix: state the shape of the distribution you want, in numbers. "80 percent of users should have 0-2 orders, 15 percent should have 3-9, 5 percent should have 10 or more" gives the agent something concrete to generate against, instead of leaving it to default to uniform. The same applies to dates, cluster more rows in recent months rather than spreading them evenly, and to prices, vary them by category instead of one flat range.

Non-idempotent scripts that fail on a second run

A seed script that only works once is a liability. It gets run again after a schema change, in CI, or by a second developer, and it either throws a duplicate key error or silently doubles every row.

The fix: name the idempotency strategy in the prompt instead of leaving it to the agent's default. For raw SQL, that means an upsert using ON CONFLICT DO NOTHING or DO UPDATE, keyed on a natural unique column like email. For an ORM like Prisma, that means upsert() instead of create() in the seed script, which is what Prisma's own seeding documentation recommends. If duplicate-safe inserts are not practical, ask for a truncate-and-reseed pattern instead: delete rows tagged as seed data, then reinsert, so reruns are explicit resets rather than accidents.

Running and checking the script before you trust it

Do not take a seed script's first successful run as proof it is correct. Run it a second time immediately, against the same database, and confirm it exits cleanly with no new duplicate rows. Then check the distribution with a quick GROUP BY on orders per user, rather than eyeballing a handful of rows. If the agent used a library like Faker.js for names and addresses, skim a sample for anything obviously wrong, like emails that do not match names or negative prices. Treat the result like any other AI-written code before it runs anywhere other than a disposable local database.

Frequently asked questions

What is the difference between a seed script and a database migration?

A migration changes the schema itself, adding or altering tables and columns. A seed script runs after the schema exists and inserts rows into it. AI coding agents that write database migrations face different failure modes, mostly around destructive changes and rollback safety, while seed scripts fail mostly around insert order, data realism, and repeatability.

How much sample data should I ask for?

Enough to exercise the behavior you are testing, not an arbitrary round number. Testing pagination needs enough rows to span several pages. Testing a dashboard aggregate needs enough variation for the number to move meaningfully. A few hundred rows per table is usually plenty for local development; load testing needs its own, much larger seed.

Should I ask the AI to invent data directly, or use a library like Faker?

Either works, but naming a library in the prompt gets more realistic output with less back and forth. Faker.js and similar libraries have pre-built generators for names, addresses, and product data that are more varied than what a model invents unassisted. Let the agent pick the library, but require varied, non-sequential output either way.

How do I make an AI-written seed script idempotent?

Say the word idempotent in the prompt and name the mechanism: upsert keyed on a unique column, ON CONFLICT DO NOTHING, or a delete-then-reinsert pattern. Then prove it by running the script twice against the same database and checking that row counts did not change on the second run.

Can AI generate realistic test data without a real schema in front of it?

Not well. A prompt asking for generic realistic test data, without column names, types, and constraints, tends to produce plausible-looking but structurally wrong output. Paste the actual schema, even if it is rough, before asking for the seed script itself.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.