Streaming LLM Responses: SSE, Buffering, and Why Output Stalls
Streaming is what makes an AI feature feel fast. It is also where proxies, buffers and framework defaults quietly break things. A practical guide.
Streaming does not make generation faster. It makes it feel faster, by showing tokens as they arrive instead of after the last one. For anything interactive that perceptual difference is the whole user experience.
How it works
Set stream: true and the response becomes a series of server-sent events rather than one JSON body:
data: {"choices":[{"delta":{"content":"Hello"}}]}
data: {"choices":[{"delta":{"content":" world"}}]}
data: [DONE]
Each chunk carries a delta — the new text only. You accumulate them. The stream ends with a sentinel, conventionally [DONE].
Three details catch people out: events are separated by a blank line, [DONE] is not valid JSON so parsing it will throw, and a single TCP read may contain a partial event. Buffer until you have a complete event before parsing.
Time to first token is the metric
Total generation time is largely fixed by output length. What you control is the delay before the first token, and that is dominated by how much input the model must process first.
This is a direct argument for keeping prompts tight: a bloated context does not just cost more, it makes the interface feel sluggish before a single character appears.
The failure that wastes the most time
Everything works locally, then in production output arrives all at once at the end. Almost always something in the path is buffering.
Usual suspects:
- Reverse proxies. Nginx buffers proxied responses by default. Disable it for streaming routes.
- Compression. Gzip middleware may wait to accumulate before compressing. Exclude event-stream responses.
- Serverless platforms. Some collect the whole response before returning. Check streaming is supported on your runtime.
- Your own framework. Middleware that reads the body to log it will consume the stream.
Diagnose from the outside in with curl -N — if tokens trickle there but not in your app, the problem is yours; if curl also stalls, it is upstream.
Usage data and streams
A non-streamed response includes a usage block. Streamed responses often do not, unless you opt in — and providers differ on how. If you bill or budget on tokens, verify you actually receive usage while streaming before you rely on it.
Cancellation matters more than it seems
When a user navigates away, abort the request. Otherwise generation continues to completion and you pay for tokens nobody will read. On a busy interface this is a meaningful, entirely avoidable cost.
Propagate the cancellation all the way upstream — aborting the browser fetch while your server keeps streaming from the provider fixes nothing.
Streaming tool calls
Tool calls stream as fragments: the function name arrives first, then arguments in pieces. Partial arguments are not valid JSON and never will be until complete. Accumulate by index, wait for the finish reason, then parse.
Rendering as it arrives
Streaming markdown creates a presentation problem: at any moment you may hold a half-written code fence or an unclosed bold marker. Naively re-rendering each chunk produces visible flicker as the parser reinterprets incomplete syntax.
Two approaches work. Render plain text until a construct completes, then upgrade it. Or use an incremental parser that tolerates unterminated blocks. Either beats re-parsing the whole buffer on every token, which also gets slow on long responses.
Errors mid-stream
A request that returns 200 and starts streaming can still fail halfway — an upstream timeout, a content filter, a dropped connection. Your client has already shown the user partial output.
Decide the behaviour deliberately: keep what arrived and mark it incomplete, or discard and retry. Silently leaving a truncated answer on screen is the one option that is always wrong, because the user cannot tell it is truncated.
Reconnection and resumption
Mobile clients drop connections routinely. SSE has a built-in reconnection mechanism via the Last-Event-ID header, but it only helps if your server can resume from an event ID — and most LLM proxies cannot, because generation is not replayable.
In practice the pragmatic choice is to treat a dropped stream as a failed request, keep the partial output visible, and let the user retry explicitly rather than silently regenerating and charging twice.
A checklist before shipping
- Confirm tokens arrive incrementally through your full production path, not just locally.
- Handle mid-stream disconnects — the model may stop partway.
- Abort upstream when the client disconnects.
- Verify you get usage data if you need it.
- Test with a slow network; that is where buffering bugs surface.
Common questions
Why does my streamed response arrive all at once?
Something in the path is buffering — commonly an Nginx proxy, gzip middleware, or a serverless runtime that collects the full response. Test with curl -N to find which side is responsible.
Does streaming reduce the cost of a request?
No. You pay for the same tokens. It improves perceived latency, and it lets you cancel early, which does save money on abandoned requests.
Can I get token usage while streaming?
Often only if you opt in, and support varies between providers. Verify it explicitly before relying on streamed usage data for billing or budgets.