Dashboard

How to Add Image Generation to an AI-Built App

A coding agent will wire up an image endpoint in about a minute. What it will not do is make the request asynchronous, cap what a user can spend, moderate the input, or plan where a few thousand images are going to live.

Steve Jefferson
Steve Jefferson
Developer Advocate
21 September 20261 min read

Image generation is unlike almost every other feature you will add to an app, because a single user action costs you real money and takes ten to sixty seconds to complete. Both of those break assumptions your code already makes. Ask a coding agent to "add image generation" and you will get a route that calls a provider and returns the URL, which works beautifully in testing and then produces a timeout, a bill, or both, the first week it is live.

Four things need handling: the request is slow, the request costs money, the input is user-controlled, and the output has to live somewhere. None of them are hard individually, and all four are skipped by default. If you are earlier in the process, how to build an app with AI covers the ground before this; what follows assumes you have an app and want to add generation to it.

1. Never Generate Inside the Request

A generation call that takes thirty seconds will outlast a default proxy timeout, hold a server connection open, and leave the user staring at a spinner with no idea whether it is working. Worse, when it does time out, the generation usually completed. You paid for an image you then threw away.

The pattern is: accept the job, return immediately, do the work elsewhere, notify when done. HTTP 202 Accepted is the status code for exactly this, and using it tells any client the work is still in flight.

javascript
// POST /api/images  ->  returns instantly with a job id
app.post('/api/images', async (req, res) => {
  const job = await db.imageJobs.insert({
    userId: req.user.id,
    prompt: req.body.prompt,
    status: 'queued',
    createdAt: new Date(),
  });
  await queue.enqueue('generate-image', { jobId: job.id });
  res.status(202).json({ jobId: job.id, status: 'queued' });
});

// GET /api/images/:jobId  ->  client polls this
app.get('/api/images/:jobId', async (req, res) => {
  const job = await db.imageJobs.findOne({
    id: req.params.jobId,
    userId: req.user.id,        // never trust the id alone
  });
  if (!job) return res.status(404).end();
  res.json({ status: job.status, url: job.url, error: job.error });
});

The worker does the generation, stores the result, and flips the row to done or failed. The client polls every couple of seconds, or you push over a websocket if you already have one. If you have not set up background work yet, adding background jobs to an AI-built app covers the queue side; image generation is the feature that most often forces the issue.

Record the provider's own request id on the job row as soon as you have it. When a job dies halfway you will want to ask the provider what happened, and without their id you cannot.

2. Put a Ceiling on What a User Can Spend

This is the part that turns into a genuinely bad week. A text request costs a fraction of a cent. An image costs cents each, and a user who discovers a generate button with no limit will happily press it two hundred times. Multiply by however many users find it interesting and you have a bill nobody approved.

Rate limiting is not sufficient on its own, because rate limiting caps speed, not total. You want both:

  • A rate limit, so one user cannot queue fifty jobs in ten seconds.

  • A quota, counted in images per user per period, checked before you enqueue.

  • A global daily cap as a backstop, so a bug or an abusive signup cannot run up an unbounded bill overnight.

javascript
const DAILY_LIMIT = 20;

async function assertQuota(userId) {
  const since = new Date(Date.now() - 24 * 60 * 60 * 1000);
  const used = await db.imageJobs.count({
    userId,
    createdAt: { $gte: since },
    status: { $ne: 'failed' },   // do not bill users for your outages
  });
  if (used >= DAILY_LIMIT) {
    const err = new Error('Daily image limit reached');
    err.status = 429;
    throw err;
  }
  return { used, remaining: DAILY_LIMIT - used };
}

Return the remaining count to the client on every response so the interface can show it. A user who can see "14 of 20 left today" does not file a support ticket when the button stops working.

Count the quota at enqueue time, not on completion, or a burst of parallel requests all pass the check before any of them finish. And excluding failed jobs is a deliberate choice: charging a user their quota for your provider's 500 is the kind of small unfairness people remember.

3. Moderate the Input, Then Trust Nothing

The moment you accept a user-supplied prompt and render the result, you are publishing user-generated content. Providers apply their own filters and will reject some requests outright, which you must handle as a normal outcome rather than an exception, but their policy is not your policy and their filter is not a guarantee.

  1. Check the prompt before you spend anything. A moderation endpoint costs a fraction of a generation, so screening first is cheaper than generating and discarding.

  2. Handle a provider refusal as an expected path. Show the user a clear message. Do not log their prompt into an error channel your whole team reads.

  3. Keep the prompt, the user id and the timestamp on the job row. When something does slip through, you need to know who asked for it.

  4. Give yourself a delete path. An admin action that removes an image and marks the job blocked, available before you need it rather than written under pressure.

Be aware that the prompt field is also an injection surface if any part of your system later feeds these images or their captions back into a model. Prompt injection through an image covers that route specifically, and it is easy to build accidentally.

4. Decide Where Images Live Before You Have Thousands

Most providers return a URL that expires, often within the hour. If you store that URL in your database and call it done, every image in your app breaks quietly and you find out from a user weeks later.

Download and store the bytes yourself, as part of the same worker job that generated them:

javascript
async function persist(job, providerUrl) {
  const bytes = await fetch(providerUrl).then(r => r.arrayBuffer());
  const key = `images/${job.userId}/${job.id}.webp`;
  await storage.put(key, await toWebp(bytes), {
    contentType: 'image/webp',
    cacheControl: 'public, max-age=31536000, immutable',
  });
  await db.imageJobs.update(job.id, { status: 'done', key, url: publicUrl(key) });
}

Convert to webp on the way in, since generated PNGs are large and you will be serving these repeatedly. Serve through a CDN. And write the deletion rule now, while the decision is cheap: images for deleted accounts go, and if your product does not need permanent history, expire unsaved generations after thirty days. Storage is inexpensive until it is a five-figure line nobody budgeted for.

Decision

Cheap now

Expensive later

Where bytes live

Your bucket, from day one

Provider URLs that expired months ago

Format

webp at generation time

Re-encoding a hundred thousand PNGs

Deletion policy

A rule in the worker

A migration and a legal question

Cost attribution

A cost column on the job row

Reverse-engineering a provider invoice

Know the Unit Cost Before You Price Anything

Put the per-image cost on the job row as you write it. It is one extra column and it converts your provider bill from a mystery into a query: cost per user, per feature, per plan. Without it you are guessing, and guessing about a per-use cost is how a flat-rate plan quietly loses money on its heaviest users.

That number is the input to every pricing decision you make afterwards, which is the subject of working out what an AI feature costs per user. Collect it from the first day the feature is live rather than reconstructing it later.

For the wider picture of what running this kind of app costs once several metered features are in play, what it costs to run an AI-built app sets out the other line items. Image generation is usually the one that surprises people, because it is the first feature where a single click has a visible price.

Frequently Asked Questions

Why should image generation run in a background job?

Because generation takes ten to sixty seconds, which exceeds common proxy timeouts and ties up a server connection. A timed-out request usually still completed at the provider, so you pay for an image the user never receives.

How do I stop users running up my image generation bill?

Combine a rate limit with a per-user quota checked before you enqueue, plus a global daily cap as a backstop. Rate limiting alone caps speed, not total spend.

Do I need to store generated images myself?

Yes. Provider URLs usually expire within hours. Download the bytes in the worker, convert to webp, store them in your own bucket, and serve through a CDN.

What happens when the provider refuses a prompt?

Treat it as a normal outcome, not an error. Screen prompts with a moderation check first to avoid paying for rejected generations, and show the user a clear message when a refusal happens.

How much does image generation cost per image?

It varies by provider, model and resolution, and it changes often. Record the actual cost on each job row rather than relying on a figure from a pricing page you read once.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.