How to Add a Bulk Import Feature to an AI-Built App
The upload button is the easy part. Here is how to handle validation, duplicates, and partial failures so bulk import does not quietly corrupt your users' data.
Every app that stores a list of anything, customers, products, invoices, eventually gets asked for a bulk import. Someone has 2,000 existing customers in a spreadsheet and does not want to type them in one at a time. It looks like a small feature. The part that actually takes the work is not the upload button, it is deciding what happens when a row is wrong.
The Upload Is the Easy 20%
Accepting a CSV file, parsing it, and inserting rows into a database is genuinely straightforward, and an AI coding agent will produce a working version of this quickly. The features that separate a bulk import people trust from one that quietly corrupts data are all in how you handle imperfect input, which is most real-world spreadsheets.
Validate Before You Insert, Not During
The single highest-leverage design decision: parse and validate the entire file first, show the user a summary of what will happen, and only write to the database after they confirm. Inserting row by row as you parse means a failure halfway through leaves the database in a half-imported state that is hard to reason about and harder to explain to the user.
Validation check | Why it matters |
|---|---|
Required fields present in every row | Catches incomplete rows before they become incomplete database records |
Data types match (a date column that is actually a date, not text) | Prevents silent type coercion errors that surface later as bugs |
Duplicate detection against existing records | Stops the same customer being created twice from re-uploaded files |
Foreign key references exist (a product ID the import references is real) | Prevents orphaned references that break the app elsewhere |
Show a Preview, Not Just a Result
Before committing anything, show the user: how many rows will be created, how many will be updated (if you support upsert), how many will be skipped and why, with the specific error for each skipped row. "3 rows skipped: missing email address" is actionable. A silent partial import discovered a week later when someone notices missing customers is a trust-destroying bug.
Prompt an AI coding agent for this explicitly: "parse the whole file first, build a summary of rows to create, rows to update, and rows with errors including the specific error message per row, and show that summary to the user before any database write happens." Left to its own judgement, an agent will often build the simpler insert-as-you-parse version, since it is a shorter code path and passes a happy-path test.
Decide Your Duplicate Policy Up Front
Skip duplicates: if a row matches an existing record on your chosen unique key (usually email or an external ID), leave the existing record untouched.
Update duplicates: overwrite the existing record's fields with the imported values, useful for periodic re-syncs from an external source.
Flag duplicates for manual review: safest default when you are not sure which of the above the user wants, and it costs little to ask them to confirm.
Whichever you choose, make it visible in the preview screen, not a silent default buried in your code. Users who do not know which policy is active will misinterpret the results either way.
Handle the File Itself Defensively
Cap the file size and row count, and reject oversized files with a clear message rather than letting a 200,000-row upload hang your import job or time out silently.
Handle encoding issues (a CSV exported from an older version of Excel is a common source of garbled characters) by detecting and normalizing encoding before parsing, not assuming UTF-8.
Do the parsing and validation in a background job for anything beyond a few hundred rows, so a large file does not tie up a web request until it times out.
Give Users a Template, Not a Blank Slate
Provide a downloadable CSV template with your exact expected column headers, and match against those headers case-insensitively with reasonable tolerance for common variations ("Email" vs "email" vs "Email Address"). Asking an AI coding agent to add fuzzy column-name matching against a known set of expected fields is a small addition that removes a large share of avoidable import failures.
a complete guide to building an app with AIadding CSV export to an AI-built appadding role-based access control to an AI-built appadding CSV import to an AI-built app
FAQ
Should bulk import run synchronously or as a background job?
For anything beyond a couple hundred rows, run it as a background job and show progress, rather than holding an HTTP request open. This also avoids the import silently failing if a user closes the browser tab mid-upload.
What file formats should I support beyond CSV?
CSV covers the vast majority of real-world use cases and is the simplest to validate reliably. Adding Excel (.xlsx) support is reasonable once you have real user demand, but it adds meaningfully more parsing complexity (multiple sheets, merged cells, formulas) for a format most spreadsheet tools can also export as CSV.
How do I handle an import that partially succeeds?
With the validate-first approach described above, partial success should be rare and intentional, rows with errors are skipped and reported, valid rows are committed. If you must support partial commits for other reasons, make the report of what succeeded and what did not exhaustive and exportable, so the user can fix and re-upload just the failed rows.
Is it safe to let AI generate the whole import feature unsupervised?
Generate it, then specifically review the duplicate-handling logic and the failure path (what happens on a malformed row) by hand. These are the two places where a plausible-looking implementation most often has a real bug, since happy-path tests do not exercise them.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


