Common API errors and how to handle them
A practical guide to handling common API errors like authentication failures, 429 rate limits, and model not found, with retry and backoff best practices.
Introduction
When integrating with an LLM API aggregator, you'll occasionally encounter errors. This guide covers common error types and how to handle them gracefully, ensuring your application remains robust and user-friendly.
Authentication Failures
Authentication errors occur when the API key is missing, invalid, or lacks permission. The API typically returns a 401 Unauthorized or 403 Forbidden status.
How to handle:
- Verify the API key is correctly included in the request headers (e.g.,
Authorization: Bearer <key>). - Check that the key is active and has sufficient balance if the service uses prepaid credits.
- Avoid hardcoding keys; use environment variables or a secure secret manager.
- If using multiple keys, ensure the correct one is selected for the environment.
Rate Limiting (429)
Rate limits protect the service from abuse and ensure fair usage. When exceeded, the API returns 429 Too Many Requests, often with a Retry-After header indicating how long to wait.
How to handle:
- Implement exponential backoff with jitter: wait for a short interval, then double it on each retry, adding randomness to avoid thundering herd.
- Respect the
Retry-Afterheader if provided. - Limit the number of retries to avoid infinite loops.
- Consider queuing requests or reducing concurrency if rate limits are frequently hit.
Model Not Found
If you request a model that doesn't exist or isn't available, the API returns 404 Not Found or a similar error. This can happen due to typos or deprecated models.
How to handle:
- Double-check the model identifier against the provider's documentation.
- Implement a fallback to a default model if the requested one is unavailable.
- Log the error for monitoring and alerting.
Retry and Backoff Best Practices
Not all errors should be retried. Distinguish between transient errors (e.g., 429, 5xx) and permanent errors (e.g., 401, 404).
Best practices:
- Retry only on transient errors.
- Use exponential backoff with jitter.
- Set a maximum retry limit (e.g., 3-5 attempts).
- Log retries for debugging.
- Consider circuit breakers to stop retrying when the service is consistently failing.
Conclusion
Handling API errors gracefully is essential for reliable applications. By implementing proper authentication, respecting rate limits, validating model names, and using smart retry strategies, you can minimize disruptions and provide a better experience for your users.