Smart routing: set model=auto
Setting model=auto lets the platform choose a suitable model for each request based on availability and task characteristics. This guide explains how auto-routing works conceptually, the tradeoffs between routing and pinning, and when to explicitly specify a model.
What model=auto does
When you set model to auto in your API request, you delegate the choice of model to the platform. Instead of you specifying a particular model, the platform selects one that is appropriate for the request. The selection is based on factors such as the task type, the input size, and current model availability. The goal is to provide a good default without requiring you to know which model is best for every situation.
How routing works conceptually
The router evaluates the request and maps it to a model. It does not expose the exact decision logic, but the process generally considers:
- Task characteristics: For example, a short conversational query versus a long document analysis.
- Model capabilities: Different models excel at different things—some are faster, some handle longer contexts, some are better at code.
- Availability and load: The router can avoid models that are temporarily overloaded or unavailable.
The router aims to balance quality, speed, and cost. Because you pay the official price multiplied by 1.3, the cost of an auto-routed request depends on which model is chosen. You are not charged extra for the routing itself.
Tradeoffs of using auto
Advantages:
- Simplicity: You don't need to track which model is best for each use case.
- Resilience: If a model is down, the router can use another.
- Potential cost savings: The router may select a cheaper model when it deems it sufficient.
Disadvantages:
- Less control: You cannot predict exactly which model will handle a request.
- Variable output style: Different models have different tones and formats, which can matter for consistency.
- Latency variance: Routing to a different model may change response times.
When to pin a specific model
Pinning means setting model to a specific identifier (e.g., claude-3-5-sonnet, gpt-4o, deepseek-chat). You should pin when:
- You need consistent behavior: For example, when building a product that requires a uniform voice or format.
- You rely on specific capabilities: Such as a particular context length, function-calling support, or coding ability.
- You have compliance or reproducibility requirements: You need to know exactly which model generated the output.
- You want predictable costs: The price per token is fixed for a given model, so you can estimate expenses more accurately.
How to use model=auto
In your API request, set the model field to "auto". For example:
{
"model": "auto",
"messages": [
{"role": "user", "content": "Summarize this article."}
]
}
The platform will route the request to a suitable model. The response will include the actual model used, so you can inspect it if needed. If you want to force a specific model, replace "auto" with the model's identifier.
Best practices
- Start with
autofor prototyping and general use cases. - Pin a model when you need deterministic behavior or specific features.
- Monitor the models returned by auto-routing to understand typical choices.
- Use auto-routing as a fallback in production: if a pinned model fails, retry with
auto.
Summary
model=auto offers a convenient way to let the platform pick a model, balancing quality, speed, and cost. It's ideal when you don't have strict requirements. Pinning a model gives you control and predictability at the expense of manual selection. Choose based on your need for consistency, specific capabilities, and cost transparency.