Streaming responses with SSE
Learn how to enable streaming responses from LLM APIs using Server-Sent Events (SSE). This guide explains the stream=true parameter, the structure of SSE chunks, and provides a minimal framework-agnostic code example for consuming streams in real time.
What is streaming and why use it?
When you send a request to an LLM API, you can choose to receive the full response all at once or as a stream of partial results. Streaming uses Server-Sent Events (SSE) to push chunks of data as they are generated. This is useful for:
- Reducing perceived latency: users see the first words almost immediately.
- Building interactive chat interfaces that feel responsive.
- Handling long responses without waiting for the entire completion.
Most LLM APIs support streaming by adding a simple parameter to the request.
Enabling streaming with stream=true
In a typical chat completion request, set stream: true in the JSON body. For example:
{
"model": "claude",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}
The server will respond with a text/event-stream instead of a single JSON object. Each event contains a chunk of the generated text.
Understanding SSE chunks
SSE responses consist of lines prefixed with data: followed by a JSON payload. A typical chunk looks like:
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","choices":[{"delta":{"content":"Hello"},"index":0}]}
Key points:
- The
deltafield holds the incremental text. You append eachdelta.contentto build the full message. - A final chunk often has
finish_reasonset, indicating the stream is complete. - Some APIs send a
data: [DONE]message to signal the end. - Chunks may arrive out of order? No, they are ordered.
Minimal framework-agnostic example
Below is a bare-bones JavaScript example using fetch and the ReadableStream API. It works in modern browsers and Node.js (v18+).
async function streamCompletion(prompt) {
const response = await fetch('https://api.example.com/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': 'Bearer YOUR_API_KEY'
},
body: JSON.stringify({
model: 'claude',
messages: [{ role: 'user', content: prompt }],
stream: true
})
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop(); // keep incomplete line
for (const line of lines) {
if (line.startsWith('data: ')) {
const data = line.slice(6);
if (data === '[DONE]') return;
try {
const json = JSON.parse(data);
const content = json.choices[0]?.delta?.content || '';
process.stdout.write(content);
} catch (e) {
// ignore malformed chunks
}
}
}
}
}
streamCompletion('Tell me a joke');
This example reads the stream chunk by chunk, decodes the bytes, splits by newline, and parses each data: line. It handles the [DONE] sentinel and prints the cumulative text.
Handling errors and edge cases
- Always check
response.okbefore reading the stream. - Some APIs include error events within the stream. Look for an
errorfield in the JSON. - Network interruptions can truncate the stream; consider implementing retries.
- For browser use, be aware of CORS if calling the API directly.
Using streaming with an aggregator
If you use an API aggregator that supports multiple models, streaming works the same way as long as the provider supports it. The aggregator forwards the SSE stream from the upstream provider. You still set stream: true and consume chunks identically.
Remember that streaming may have slight differences in chunk format across providers (e.g., field names like delta vs text). Always check the API documentation for the specific model you are using.
Conclusion
Streaming with SSE is a straightforward way to improve user experience by delivering LLM responses incrementally. With the stream=true parameter and a simple reader loop, you can build real-time applications without complex infrastructure.