Mid-Stream Errors: What to Do When SSE Cuts Off After Tokens Started
Learn how to detect and handle mid-stream errors in SSE when tokens have already started arriving. This guide covers practical strategies for partial output recovery and robust error handling in LLM API streaming.
Why Mid-Stream Errors Happen
Server-Sent Events (SSE) deliver tokens incrementally, but the connection can break after some content has already been sent. Common causes include network timeouts, server-side errors, or client disconnects. Unlike errors before the first token, mid-stream errors leave you with partial output that may still be useful.
Detecting Mid-Stream Errors
The key challenge is distinguishing between a clean stream end and an abrupt cutoff. SSE streams typically end with a [DONE] sentinel or close the connection. If the connection closes or an error event arrives without [DONE], you have a mid-stream error.
- Monitor the event stream for error events: Some APIs send an
event: errorbefore closing. - Track whether
[DONE]was received: If the stream ends without it, treat it as incomplete. - Watch for unexpected disconnects: Network errors or timeouts will throw exceptions in your HTTP client.
Recovering Partial Output
When an error occurs mid-stream, you already have some tokens. Preserve and use them:
- Accumulate tokens as they arrive in a buffer. This ensures you have the partial output even if the stream fails.
- On error, decide whether to retry or use partial data. For some use cases, partial output is sufficient. For others, you may want to retry the request.
- If retrying, consider using a fresh request rather than resuming, unless the API supports continuation. Most LLM APIs do not support resuming a stream.
Implementing a Robust SSE Client
A resilient SSE client should handle mid-stream errors gracefully. Here's a conceptual approach:
- Use a try/except block around the stream reading loop.
- Maintain a buffer to accumulate the response.
- On exception, check if buffer is non-empty. If so, you have partial output.
- Log the error and the partial output for debugging.
- Decide on retry logic based on your application's needs.
When to Retry vs. Use Partial Output
Not all mid-stream errors require a retry. Consider these factors:
- Criticality of completeness: If the full response is essential (e.g., code generation), retry. If a partial answer is acceptable (e.g., chat), you might use what you have.
- Cost and latency: Retrying consumes additional tokens and time. Balance this against the value of a complete response.
- Idempotency: Ensure retries do not cause duplicate side effects.
Best Practices for Streaming Resilience
- Set reasonable timeouts on your HTTP client to avoid hanging indefinitely.
- Implement exponential backoff for retries to avoid overwhelming the server.
- Use a unique request ID if the API supports it, to help with debugging and deduplication.
- Consider fallback models: If a particular model consistently fails mid-stream, you might switch to another model available through your API aggregator.
Conclusion
Mid-stream errors are a fact of life with SSE. By detecting them early and handling partial output intelligently, you can build more resilient applications. Always accumulate tokens, decide on retry strategies based on context, and implement robust error handling in your SSE client.