Dashboard

How to Seed Test Data in an AI-Built App

Generated sample data passes every test and hides every real bug. How to write a seed script that reproduces the shapes that actually break an app.

Steve Jefferson
Steve Jefferson
Developer Advocate
19 September 20261 min read

How to Seed Test Data in an AI-Built App

Ask an AI builder for sample data and you get twelve users called John Smith with sequential emails, all created five seconds apart. That data passes every test you write against it and hides every bug you actually have. A seed script is worth building deliberately, and it takes about twenty minutes.

What good seed data has that generated filler does not

The point of seeding is to reproduce the shapes that break things. Four properties do most of the work:

Property

Why it catches bugs

Spread over real time

Sorting, pagination and "recent" queries only break when timestamps are not sequential

Awkward strings

Apostrophes, accents, emoji in a display name, and a 200-character company name break layouts and escaping

Empty and maximal states

A user with zero orders and a user with 400 are the two screens nobody designs

Referential edges

An order whose customer was deleted, a record with a null optional field

Filler data has none of these, which is why an app can look fine in development and fall over on the first real signup with an apostrophe in their surname.

Write the script as code, not as a prompt

The important rule is that seeding must be repeatable and checked into the repository. Asking an AI to "add some test data" writes rows directly into your database and leaves no record of what it did, so the next person cannot reproduce your state and neither can CI. Ask for a script instead.

typescript
// seed.ts  -- run with: npm run seed
import { faker } from '@faker-js/faker'
import { db } from './src/db'

faker.seed(42)                     // same data every run, so tests are stable

const EDGE_NAMES = [
  "O'Brien", "Zoe", "Muller", "Nakamura",
  "A Very Long Company Name That Will Wrap On Every Card Component Ltd",
]

async function main() {
  await db.order.deleteMany({})    // idempotent: safe to re-run
  await db.user.deleteMany({})

  // 1. the two states nobody designs
  await db.user.create({ data: { name: 'Zero Orders', email: 'zero@example.test' } })
  const heavy = await db.user.create({ data: { name: 'Four Hundred Orders', email: 'heavy@example.test' } })
  for (let i = 0; i < 400; i++) {
    await db.order.create({ data: { userId: heavy.id, total: 100, createdAt: daysAgo(i) } })
  }

  // 2. awkward strings
  for (const name of EDGE_NAMES) {
    await db.user.create({ data: { name, email: faker.internet.email() } })
  }

  // 3. the ordinary middle, spread over a year
  for (let i = 0; i < 50; i++) {
    await db.user.create({
      data: {
        name: faker.person.fullName(),
        email: faker.internet.email(),
        createdAt: daysAgo(faker.number.int({ min: 0, max: 365 })),
      },
    })
  }
}

const daysAgo = (n: number) => new Date(Date.now() - n * 86_400_000)

main().then(() => process.exit(0))

Two lines there matter more than the rest. faker.seed(42) makes the output identical on every run, which is what lets a test assert on a specific record. The deleteMany calls at the top make the script idempotent, so running it twice does not give you a hundred users.

Use a reserved domain for fake addresses

Seed addresses should use example.test or example.com, both reserved by standards for exactly this. Faker's default domains are real, and a staging environment that accidentally has email sending switched on will cheerfully mail strangers. This is a genuinely common incident, and it costs nothing to avoid.

Keep seeds out of production

Guard the script so it cannot run against a production connection string. One check at the top is enough, and it has saved more data than any backup policy:

typescript
if (process.env.NODE_ENV === 'production' || /prod/i.test(process.env.DATABASE_URL ?? '')) {
  throw new Error('refusing to seed: this looks like production')
}

Where seeding fits with the rest of your testing

Seed data is the input your tests run against, so it is worth getting right before you write many of them. If you are asking an AI to generate the test suite as well, give it the seed script as context first, otherwise it will invent its own fixtures and you end up with two sources of truth. There is more on that in getting AI to write the tests themselves, and the broader pass is in testing the app properly before launch.

Some of this depends on what you are running underneath: a document store and a relational database have different answers for the referential-edge cases above. If that is still open, see picking the database in the first place. The wider context sits in the end-to-end build guide.

Frequently asked questions

How much seed data do I need?

Enough to make pagination real, so more than one page, and enough to include your edge cases. Fifty to a hundred ordinary records plus a handful of deliberate outliers covers most apps.

Should seed data be random or fixed?

Fixed, via a seeded random generator. Random data makes tests flaky and makes failures impossible to reproduce.

Can I use a copy of production data instead?

Only if you anonymise it properly, and anonymising it properly is harder than writing a seed script. Real data in a development environment is a breach waiting for a misconfiguration.

Where should the seed script live?

In the repository, next to your migrations, with a npm script that runs it. If it is not in version control it does not exist for anyone but you.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.