sse-streaming-python-node · EN · 2026-10-08

Handling SSE Streams in Python and Node Without Leaking Connections

Learn how to handle Server-Sent Events (SSE) streams from LLM APIs in Python and Node.js without leaking connections. Compare patterns using OpenAI SDKs and raw HTTP, and discover best practices for cleanly closing sockets when streams end early.

Understanding SSE Streams

Server-Sent Events (SSE) are a standard for streaming text over HTTP. LLM APIs use SSE to send tokens as they are generated. A typical SSE response consists of lines like:

data: {"choices":[{"delta":{"content":"Hello"}}]}

Each event is separated by a blank line. The stream ends when the server sends a data: [DONE] message and closes the connection.

Why Connection Leaks Happen

If a stream is not fully consumed or the underlying socket is not closed, the connection may remain open. This can lead to:

  • Resource exhaustion on the client.
  • Hitting rate limits or connection caps.
  • Memory leaks in long-running processes.

Common causes include:

  • Breaking out of a loop reading the stream without closing the response.
  • Not handling exceptions that occur mid-stream.
  • Using a library that buffers the entire response instead of streaming.

Using OpenAI SDKs (Python & Node)

OpenAI SDKs provide built-in streaming support. They manage the HTTP connection for you, but you still need to handle early termination.

Python

from openai import OpenAI

client = OpenAI(api_key="your-key")

stream = client.chat.completions.create(
    model="gpt-4",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True
)

try:
    for chunk in stream:
        content = chunk.choices[0].delta.content
        if content:
            print(content, end="")
        # Break early if needed
        # if some_condition:
        #     break
except Exception as e:
    print(f"Error: {e}")
finally:
    stream.close()  # Ensure the connection is closed

Key points:

  • The stream object is an iterator. When you break early, call stream.close() to release the connection.
  • Use try/finally to guarantee closure even if an exception occurs.

Node.js

import OpenAI from 'openai';

const client = new OpenAI({ apiKey: 'your-key' });

async function main() {
  const stream = await client.chat.completions.create({
    model: 'gpt-4',
    messages: [{ role: 'user', content: 'Hello' }],
    stream: true,
  });

  try {
    for await (const chunk of stream) {
      const content = chunk.choices[0]?.delta?.content;
      if (content) process.stdout.write(content);
      // Break early if needed
      // if (someCondition) break;
    }
  } catch (err) {
    console.error('Error:', err);
  } finally {
    // The SDK automatically closes the connection when the stream is fully consumed or an error occurs.
    // However, if you break early, you should call stream.controller.abort() to close the connection.
    stream.controller.abort();
  }
}

main();

Key points:

  • The stream is an async iterable. Breaking out of the loop does not automatically close the connection.
  • Use stream.controller.abort() to cancel the request and close the connection.

Raw HTTP with SSE

Sometimes you may need to use raw HTTP to have more control. Here’s how to handle SSE streams safely.

Python with requests

import requests
import json

url = "https://api.example.com/v1/chat/completions"
headers = {"Authorization": "Bearer your-key"}
data = {
    "model": "gpt-4",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": True
}

with requests.post(url, headers=headers, json=data, stream=True) as response:
    response.raise_for_status()
    for line in response.iter_lines():
        if line:
            line = line.decode('utf-8')
            if line.startswith('data: '):
                data_str = line[6:]
                if data_str == '[DONE]':
                    break
                try:
                    chunk = json.loads(data_str)
                    content = chunk['choices'][0]['delta'].get('content')
                    if content:
                        print(content, end='')
                except json.JSONDecodeError:
                    pass
    # The context manager ensures the response is closed.

Key points:

  • Use stream=True to get a streaming response.
  • The with statement ensures the connection is closed even if an exception occurs or you break early.

Node.js with fetch

const url = 'https://api.example.com/v1/chat/completions';
const headers = {
  'Authorization': 'Bearer your-key',
  'Content-Type': 'application/json',
};
const body = JSON.stringify({
  model: 'gpt-4',
  messages: [{ role: 'user', content: 'Hello' }],
  stream: true,
});

async function main() {
  const controller = new AbortController();
  const response = await fetch(url, {
    method: 'POST',
    headers,
    body,
    signal: controller.signal,
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';

  try {
    while (true) {
      const { done, value } = await reader.read();
      if (done) break;

      buffer += decoder.decode(value, { stream: true });
      const lines = buffer.split('\n');
      buffer = lines.pop(); // Keep incomplete line

      for (const line of lines) {
        if (line.startsWith('data: ')) {
          const data = line.slice(6);
          if (data === '[DONE]') {
            controller.abort(); // Close the connection
            return;
          }
          try {
            const chunk = JSON.parse(data);
            const content = chunk.choices[0]?.delta?.content;
            if (content) process.stdout.write(content);
          } catch (e) {
            // Ignore JSON parse errors for incomplete chunks
          }
        }
      }
    }
  } catch (err) {
    if (err.name !== 'AbortError') {
      console.error('Error:', err);
    }
  } finally {
    reader.releaseLock();
    controller.abort(); // Ensure connection is closed
  }
}

main();

Key points:

  • Use an AbortController to cancel the request when done or on error.
  • Always call reader.releaseLock() and controller.abort() to release resources.

Best Practices to Avoid Leaks

  • Always close the stream or response in a finally block.
  • For SDKs, check the documentation for the correct way to cancel a stream (e.g., close() in Python, abort() in Node).
  • When breaking early, ensure you close the connection.
  • Handle exceptions gracefully and still close the connection.
  • Avoid buffering the entire response in memory; process chunks as they arrive.
  • Set a timeout for the request to avoid hanging connections.

Conclusion

Handling SSE streams correctly is crucial for reliable and efficient LLM API usage. Whether you use the OpenAI SDKs or raw HTTP, always ensure connections are closed when the stream ends early. This prevents resource leaks and keeps your application responsive.

If you're looking for a hassle-free way to access multiple LLMs with a single API key, consider using our service. Top up with USDC on Base, no KYC required. You pay official price × 1.3, and key contributors can earn credits. Check our docs for more details.