auto-routing-cost-cap · EN · 2026-10-09

Capping Spend with model=auto: Setting a Cost Ceiling Instead of Naming a Model

Instead of hardcoding a model name, you can set a cost ceiling per request by using `model=auto` with a budget parameter. This article explains how auto routes requests based on your budget, how the 1.3x charge applies, and how to reconcile costs afterward.

Why Set a Cost Ceiling?

Normally, you pick a specific model (e.g., claude-3-sonnet or gpt-4) and pay its per-token price. But sometimes you care more about the budget than the model itself. For example, you might want to spend at most $0.02 on a request and are happy with any model that fits. With model=auto, you can set a cost ceiling and let the API choose a model that stays within it.

How auto Works with a Budget

When you call the API with model=auto and a budget parameter (e.g., max_cost=0.02), the system does the following:

  • Estimates the cost of the request based on the prompt length and expected output length.
  • Filters the available models to those whose estimated cost is at or below your budget.
  • Selects the highest-quality model among those that fit (or a balanced choice, depending on configuration).
  • Routes your request to that model.

If no model fits within the budget, the request fails with an error indicating that the budget is too low.

Across Cheap and Expensive Tiers

You can set different budgets to steer toward different tiers:

  • Very low budget: Routes to economical models (e.g., small or distilled variants).
  • Moderate budget: May select mid-tier models that balance cost and capability.
  • High budget: Can include premium models (e.g., Claude, GPT-4 class).

The actual model chosen depends on your input size and the current pricing of each model. For long prompts, even cheap models might exceed a low budget, so the system might reject the request.

Reconciliation and the 1.3x Charge

All usage through the aggregator is charged at the official price times 1.3. This means your actual cost will be 30% higher than the model's list price. When you set a budget, the system uses the 1.3x rate internally to estimate costs. So max_cost=0.02 means you are willing to pay up to $0.02 total, including the markup.

After the request, you can see the exact model used and the cost in your dashboard. The cost shown already includes the 1.3x multiplier. If you are a key contributor, you receive credit at official price × 1.1 (or × 1.2 for premium) in USDC, which is separate from your spending.

Best Practices

  • Start with a generous budget to avoid failures, then tighten it as you learn typical costs.
  • Use model=auto for batch or non-critical tasks where any capable model works.
  • For production workloads where you need consistency, specify a model directly.
  • Monitor your usage to understand which models are selected at different budgets.

Example: Setting a Cost Ceiling

import requests

response = requests.post(
    "https://api.example.com/v1/chat/completions",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json={
        "model": "auto",
        "max_cost": 0.02,
        "messages": [{"role": "user", "content": "Summarize this article."}]
    }
)

In this example, the system will try to fulfill the request with a model that costs no more than $0.02 (including the 1.3x markup). If successful, the response will include the chosen model in the metadata.

Conclusion

Using model=auto with a cost ceiling lets you control spending without micromanaging model selection. It is especially useful when you have a fixed budget per request and are flexible about which model handles it. Just remember that the 1.3x charge applies, so your effective budget is 30% lower than the official price would suggest.