How to Stream an AI Response in an App
Streaming makes an AI app feel fast because words appear as they are generated. A minimal server and browser pattern, and the four bugs that break it in production.
How to stream an AI response in an app comes down to one pattern: have your server request a streamed completion from the model API and forward each chunk to the browser as it arrives, usually over server-sent events or a streamed fetch response. The browser reads the chunks and appends them to the screen. The model takes the same time to finish, but the user sees text after a fraction of a second instead of staring at a spinner.
The code is short. The failures are in the details around it: buffering proxies, cancelled requests, broken formatting, and duplicate sends. This guide shows the minimal pattern and then the four problems that appear once real users arrive.
Why streaming matters more than raw speed
Perceived speed is mostly about time to first visible word. A response that takes eight seconds in total feels fine if text starts at half a second, and feels broken if nothing shows for eight. The metric behind this is covered in what is time to first token.
The server: forward chunks as they arrive
Provider SDKs differ in names, but they all offer a streaming mode that yields pieces of text. Your server's job is to loop over them and write each one to the response without waiting for the whole answer. Here is the shape, with the provider call left abstract:
// Express-style route. streamFromModel() is your provider's streaming call.
app.post('/api/chat', async (req, res) => {
res.setHeader('Content-Type', 'text/event-stream');
res.setHeader('Cache-Control', 'no-cache, no-transform');
res.setHeader('Connection', 'keep-alive');
res.flushHeaders();
const controller = new AbortController();
req.on('close', () => controller.abort()); // user left, stop paying
try {
for await (const piece of streamFromModel(req.body.messages, controller.signal)) {
res.write(`data: ${JSON.stringify({ text: piece })}\n\n`);
}
res.write('data: [DONE]\n\n');
} catch (err) {
res.write(`data: ${JSON.stringify({ error: 'stream_failed' })}\n\n`);
} finally {
res.end();
}
});The browser: read the stream and append
The browser's built-in EventSource only supports GET requests, which is awkward when you need to send a conversation. A plain fetch with a stream reader handles POST fine:
const res = await fetch('/api/chat', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ messages }),
signal: abortController.signal,
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { value, done } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const events = buffer.split('\n\n');
buffer = events.pop(); // keep the incomplete tail
for (const e of events) {
if (!e.startsWith('data: ')) continue;
const payload = e.slice(6);
if (payload === '[DONE]') return;
appendToMessage(JSON.parse(payload).text);
}
}The line that keeps the incomplete tail is the one beginners miss. Network chunks do not respect event boundaries, so a single event can arrive split across two reads.
Four bugs that appear in production
Symptom | Usual cause | Fix |
|---|---|---|
Text arrives all at once at the end | A proxy or CDN buffers the response | Send no-transform and no-cache headers, and disable response buffering on your host or reverse proxy |
Model keeps running after the user leaves | No abort wiring | Abort the upstream request when the client connection closes, as in the server code above |
Bold, lists or code blocks flicker or show raw symbols | Markdown rendered from a half-finished string | Render on a throttle, and tolerate unclosed formatting until the next chunk |
Duplicated or interleaved messages | Double-click or retry sends two streams | Disable the send button while streaming and ignore a second request with the same message id |
Buffering is the most common one, and it hides in local testing because your laptop has no proxy in front. Always test streaming on the deployed environment.
Handling errors in the middle of a stream
Once you have sent a 200 status and started writing, you cannot change the status code. A failure mid-stream has to travel inside the stream, which is why the example sends an error event. On the browser side, keep the partial text visible and offer a retry rather than wiping the message.
Rate limits come earlier, before the first chunk, and are better handled with retries. See AI API rate limits explained and how to add rate limiting to an AI-built app.
Where this fits when you build with AI
If you generate your app with an AI builder, ask it for this pattern by name: server-sent events, abort on disconnect, and a throttled markdown renderer. Generated code often streams on the happy path and skips the abort and the buffering headers. The wider project context is in how to build an app with AI, and for long jobs that should not hold a connection open at all, background jobs are the better fit. Official reference for the wire format is the MDN guide to server-sent events.
FAQ
Do I need WebSockets to stream AI responses?
No. Server-sent events or a streamed fetch response are enough for one-way text from server to browser, and they work through more proxies than WebSockets.
Why does my streamed AI response show up all at once?
Something between your server and the browser is buffering. Check proxy and CDN settings, and confirm your headers disable caching and transformation.
Does streaming make the AI answer faster?
No. Total generation time is unchanged. Users see the first words sooner, which is what makes the app feel faster.
Can I stream and still save the full answer?
Yes. Collect the pieces on the server as you forward them, then store the joined text once the stream ends or fails.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


