When a Specific Vendor Model Is Down: Who to Contact and How to Escalate
When a specific vendor model is unreachable through an LLM API aggregator, knowing whether the issue lies with the aggregator or the upstream provider determines the correct escalation path. This guide explains how to diagnose the problem, what to check first, and how to escalate effectively to minimize downtime.
Introduction
When you call a model like Claude, GPT, or DeepSeek through an LLM API aggregator and get an error, the first question is: is the problem on the aggregator's side, or is the upstream provider down? The answer determines who to contact and how to escalate. This guide walks through the diagnostic steps and escalation paths for both scenarios.
Step 1: Diagnose the Issue
Before escalating, gather basic information to pinpoint the source.
- Check the error message: Aggregator-side errors often mention the aggregator's service, while upstream errors may include the provider's name or specific error codes (e.g., 429, 503).
- Test other models: If only one model fails while others work, the issue is likely with that specific upstream provider. If all models fail, the aggregator may be experiencing a broader outage.
- Try a direct call: If you have direct access to the upstream provider, test the same model directly. If it also fails, the provider is down. If it works, the aggregator may be having trouble reaching the provider.
- Check status pages: Look for official status pages from both the aggregator and the upstream provider. Many providers post real-time updates on outages.
Step 2: Understand the Difference
Aggregator-side issues vs. upstream provider issues require different escalation paths.
- Aggregator-side issues: Problems with authentication, billing, rate limiting, routing, or the aggregator's own infrastructure. These affect all models or a subset due to internal errors.
- Upstream provider issues: The provider's API is down, degraded, or rate-limiting requests. This affects only that provider's models (e.g., all Claude models, or all GPT models).
Step 3: Escalation Path for Upstream Provider Issues
If the issue is with a specific upstream provider:
- Check the provider's status page: Most major providers (Anthropic, OpenAI, DeepSeek, etc.) maintain status pages with incident reports. This is the fastest way to confirm an outage.
- Wait for resolution: Upstream outages are typically resolved by the provider. The aggregator cannot fix them, but may provide updates.
- Contact the aggregator's support: If the provider's status page shows no issues but you still can't reach the model through the aggregator, the aggregator's support team can investigate whether there's a routing or connectivity problem.
- Use fallback models: If your application supports it, route requests to a different model from another provider until the issue is resolved.
Step 4: Escalation Path for Aggregator-Side Issues
If the issue is with the aggregator itself:
- Check the aggregator's status page: Many aggregators provide a status page or dashboard showing service health.
- Contact support: Reach out via the aggregator's official support channels (email, chat, or ticket system). Provide details: your API key (if safe to share), timestamps, error messages, and the models affected.
- Escalate if unresolved: If the issue is critical and not resolved promptly, ask for an escalation path. Some aggregators offer priority support for key contributors or high-volume users.
- Monitor updates: Keep an eye on the aggregator's communication channels for updates.
Step 5: Best Practices for Minimizing Downtime
- Implement retries with backoff: Transient errors often resolve with retries.
- Use multiple providers: If your application can fall back to another model, configure it to do so.
- Monitor proactively: Set up alerts for error rates and latency to catch issues early.
- Document your escalation process: Know who to contact and what information to provide before an incident occurs.
Conclusion
When a specific vendor model is down, the first step is to determine whether the issue is with the upstream provider or the aggregator. Upstream issues require waiting for the provider to resolve, while aggregator issues should be reported to the aggregator's support team. By following the diagnostic steps and best practices above, you can minimize downtime and escalate effectively.