Handling SSE Streams in Python and Node Without Leaking Connections
Learn how to handle Server-Sent Events (SSE) streams from LLM APIs in Python and Node.js without leaking connections. Compare patterns using OpenAI SDKs and raw HTTP, and discover best practices for cleanly closing sockets when streams end early.
Understanding SSE Streams
Server-Sent Events (SSE) are a standard for streaming text over HTTP. LLM APIs use SSE to send tokens as they are generated. A typical SSE response consists of lines like:
data: {"choices":[{"delta":{"content":"Hello"}}]}
Each event is separated by a blank line. The stream ends when the server sends a data: [DONE] message and closes the connection.
Why Connection Leaks Happen
If a stream is not fully consumed or the underlying socket is not closed, the connection may remain open. This can lead to:
- Resource exhaustion on the client.
- Hitting rate limits or connection caps.
- Memory leaks in long-running processes.
Common causes include:
- Breaking out of a loop reading the stream without closing the response.
- Not handling exceptions that occur mid-stream.
- Using a library that buffers the entire response instead of streaming.
Using OpenAI SDKs (Python & Node)
OpenAI SDKs provide built-in streaming support. They manage the HTTP connection for you, but you still need to handle early termination.
Python
from openai import OpenAI
client = OpenAI(api_key="your-key")
stream = client.chat.completions.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello"}],
stream=True
)
try:
for chunk in stream:
content = chunk.choices[0].delta.content
if content:
print(content, end="")
# Break early if needed
# if some_condition:
# break
except Exception as e:
print(f"Error: {e}")
finally:
stream.close() # Ensure the connection is closed
Key points:
- The
streamobject is an iterator. When you break early, callstream.close()to release the connection. - Use
try/finallyto guarantee closure even if an exception occurs.
Node.js
import OpenAI from 'openai';
const client = new OpenAI({ apiKey: 'your-key' });
async function main() {
const stream = await client.chat.completions.create({
model: 'gpt-4',
messages: [{ role: 'user', content: 'Hello' }],
stream: true,
});
try {
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content;
if (content) process.stdout.write(content);
// Break early if needed
// if (someCondition) break;
}
} catch (err) {
console.error('Error:', err);
} finally {
// The SDK automatically closes the connection when the stream is fully consumed or an error occurs.
// However, if you break early, you should call stream.controller.abort() to close the connection.
stream.controller.abort();
}
}
main();
Key points:
- The
streamis an async iterable. Breaking out of the loop does not automatically close the connection. - Use
stream.controller.abort()to cancel the request and close the connection.
Raw HTTP with SSE
Sometimes you may need to use raw HTTP to have more control. Here’s how to handle SSE streams safely.
Python with requests
import requests
import json
url = "https://api.example.com/v1/chat/completions"
headers = {"Authorization": "Bearer your-key"}
data = {
"model": "gpt-4",
"messages": [{"role": "user", "content": "Hello"}],
"stream": True
}
with requests.post(url, headers=headers, json=data, stream=True) as response:
response.raise_for_status()
for line in response.iter_lines():
if line:
line = line.decode('utf-8')
if line.startswith('data: '):
data_str = line[6:]
if data_str == '[DONE]':
break
try:
chunk = json.loads(data_str)
content = chunk['choices'][0]['delta'].get('content')
if content:
print(content, end='')
except json.JSONDecodeError:
pass
# The context manager ensures the response is closed.
Key points:
- Use
stream=Trueto get a streaming response. - The
withstatement ensures the connection is closed even if an exception occurs or you break early.
Node.js with fetch
const url = 'https://api.example.com/v1/chat/completions';
const headers = {
'Authorization': 'Bearer your-key',
'Content-Type': 'application/json',
};
const body = JSON.stringify({
model: 'gpt-4',
messages: [{ role: 'user', content: 'Hello' }],
stream: true,
});
async function main() {
const controller = new AbortController();
const response = await fetch(url, {
method: 'POST',
headers,
body,
signal: controller.signal,
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
try {
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop(); // Keep incomplete line
for (const line of lines) {
if (line.startsWith('data: ')) {
const data = line.slice(6);
if (data === '[DONE]') {
controller.abort(); // Close the connection
return;
}
try {
const chunk = JSON.parse(data);
const content = chunk.choices[0]?.delta?.content;
if (content) process.stdout.write(content);
} catch (e) {
// Ignore JSON parse errors for incomplete chunks
}
}
}
}
} catch (err) {
if (err.name !== 'AbortError') {
console.error('Error:', err);
}
} finally {
reader.releaseLock();
controller.abort(); // Ensure connection is closed
}
}
main();
Key points:
- Use an
AbortControllerto cancel the request when done or on error. - Always call
reader.releaseLock()andcontroller.abort()to release resources.
Best Practices to Avoid Leaks
- Always close the stream or response in a
finallyblock. - For SDKs, check the documentation for the correct way to cancel a stream (e.g.,
close()in Python,abort()in Node). - When breaking early, ensure you close the connection.
- Handle exceptions gracefully and still close the connection.
- Avoid buffering the entire response in memory; process chunks as they arrive.
- Set a timeout for the request to avoid hanging connections.
Conclusion
Handling SSE streams correctly is crucial for reliable and efficient LLM API usage. Whether you use the OpenAI SDKs or raw HTTP, always ensure connections are closed when the stream ends early. This prevents resource leaks and keeps your application responsive.
If you're looking for a hassle-free way to access multiple LLMs with a single API key, consider using our service. Top up with USDC on Base, no KYC required. You pay official price × 1.3, and key contributors can earn credits. Check our docs for more details.