Configuring Your Client So a Key Problem Never Takes Down Your Whole App
Learn how to configure API clients with timeouts, retries, and fallback routing so that a single key issue—like invalid credentials, quota exhaustion, or provider outages—doesn't crash your entire application. This guide covers practical patterns for resilient LLM API usage.
Why Key Problems Shouldn't Be Fatal
When your app calls an LLM API, a single key issue—expired credit, revoked access, or regional outage—can cascade into a full outage if your client treats every error the same. The fix is to design your client layer so that key-level failures are contained, not fatal.
Common Key Problems and Their Symptoms
- Invalid or expired key: Immediate
401 Unauthorizedresponses. - Quota exhausted or billing issue:
402 Payment Requiredor429 Too Many Requestswith a billing hint. - Provider-side outage:
5xxerrors or timeouts. - Network or DNS failure: Connection errors before any HTTP response.
Each of these should trigger a different response from your client.
Configuring Timeouts and Retries
Set conservative timeouts to avoid hanging requests. Use exponential backoff with jitter for retries, but only for transient errors (e.g., 429, 5xx). Do not retry on 401 or 402—those require human intervention or a fallback key.
import httpx
from tenacity import retry, stop_after_attempt, wait_exponential_jitter, retry_if_exception_type
@retry(
stop=stop_after_attempt(3),
wait=wait_exponential_jitter(initial=1, max=10),
retry=retry_if_exception_type((httpx.TimeoutException, httpx.HTTPStatusError)),
)
def call_llm(payload):
# only retry on 5xx or 429
response = client.post(url, json=payload, timeout=30)
if response.status_code in (401, 402):
raise NonRetryableError(response)
response.raise_for_status()
return response.json()
Implementing Fallback Routing
If your primary key fails, route to a secondary key or provider. Many aggregators (like this one) let you manage multiple keys or use a single key across providers. Design your client to switch on specific error classes:
- On
401/402: try a backup key from a different account or provider. - On
5xx: try the same request on a different model or provider.
Keep fallbacks ordered by priority and cost. For example, if you use this platform, you might fall back from a premium model to a cheaper one, knowing you pay official price × 1.3.
Circuit Breaker Pattern
Prevent repeated calls to a failing key by implementing a circuit breaker. After N consecutive failures, open the circuit for a cooldown period, during which requests are immediately routed to fallbacks. This avoids wasting time on a dead key and reduces load on your fallback.
Monitoring and Alerts
Track error rates per key and per provider. Alert on spikes in 401/402 (key issues) versus 5xx (provider issues). Log the key identifier (hashed) and error type to debug quickly. With this platform, you can also monitor your USDC balance and top up proactively.
Example: Minimal Resilient Client
class ResilientLLMClient:
def __init__(self, keys):
self.keys = keys # list of (key, base_url) tuples
def call(self, payload):
for key, base_url in self.keys:
try:
return self._call_with_retries(key, base_url, payload)
except NonRetryableError as e:
if e.status_code in (401, 402):
continue # try next key
raise
except Exception:
continue # network or 5xx after retries
raise AllKeysFailedError()
Best Practices
- Treat key problems as transient and routable, not fatal.
- Differentiate error types: retry on transient, fallback on key issues.
- Use multiple keys or providers to avoid single points of failure.
- Monitor and alert on key health.
- Keep fallback logic configurable, not hardcoded.
By isolating key failures and adding fallback layers, you ensure that one bad key doesn't take down your whole app.