Billing Disputes: When Charged More Than Expected on a 1.3x Aggregator
Learn how to reconcile token usage with the 1.3x billing multiplier on an LLM API aggregator, identify discrepancies, and file an effective support ticket.
Why Your Bill Might Look Higher Than Expected
When you use an LLM API aggregator that charges a 1.3x multiplier on official model prices, your invoice reflects the base model cost plus the aggregator's service fee. It's common to see a total that seems higher than your own back-of-the-envelope calculation. Before assuming an error, it helps to understand how token usage is counted and billed.
How the 1.3x Billing Rule Works
- Official price is the per-token rate published by the model provider (e.g., Claude, GPT, DeepSeek).
- Your cost is calculated as:
official price × 1.3. - The 1.3x factor covers payment processing, infrastructure, and the convenience of a single API key across multiple models.
This multiplier applies uniformly to both input and output tokens, and to all supported models. There are no hidden tiers or volume discounts unless explicitly stated.
Common Causes of Billing Discrepancies
Before filing a dispute, check these frequent sources of confusion:
- Token counting differences: Your tokenizer may count tokens differently than the provider's official tokenizer, especially for non-English text or code.
- Streaming vs. non-streaming: Some APIs return usage data only for non-streaming responses; if you stream and count manually, you might miss the final usage chunk.
- Cached tokens: If the provider offers prompt caching, cached input tokens are often billed at a different rate. Verify how your aggregator handles these.
- Failed or retried requests: Network errors or retries can generate billable tokens even if you didn't receive a successful response.
- Model version drift: Providers occasionally update tokenization or pricing. Ensure you're comparing against the current official price.
- Rounding and minimum charges: Some providers round up token counts or apply minimum charges per request. Your aggregator may pass these through.
Step-by-Step: Reconcile Your Usage
- Export your usage logs from the aggregator dashboard. Look for per-request records with input tokens, output tokens, model name, and timestamp.
- Obtain the official price list for each model you used, from the provider's public pricing page. Confirm the price at the time of usage.
- Calculate expected cost for each request:
(input tokens × official input price + output tokens × official output price) × 1.3. - Sum expected costs and compare to your actual charges.
- Identify outliers: If the difference is more than a few percent, investigate specific requests that contributed most to the gap.
- Check for non-token fees: Some aggregators charge for features like key management or priority routing. These should be listed separately.
When to File a Support Ticket
If after reconciliation you still see a material discrepancy—for example, you're charged for requests you never made, or the multiplier appears to be applied incorrectly—it's time to contact support.
What to Include in Your Ticket
- Account ID and the affected API key (masked).
- Date range of the disputed charges.
- Request IDs or timestamps for the specific transactions in question.
- Your calculated expected cost and the actual charged amount, with a clear breakdown.
- Supporting evidence: screenshots of your logs, the official pricing page, and any relevant code snippets that show how you count tokens.
How to Write a Clear Ticket
- Be concise. State the problem, the evidence, and what resolution you seek.
- Avoid emotional language. Stick to facts and numbers.
- If you've identified a pattern (e.g., all discrepancies involve a particular model), mention it.
What to Expect After Filing
Support teams typically respond within a few business days. They may ask for additional details or run their own reconciliation. If the error is on their side, they may issue a credit in USDC to your account. If the discrepancy is due to token counting differences, they'll explain the methodology.
Remember that the 1.3x multiplier is applied to official prices, so your bill will always be higher than the raw provider cost. The goal of a dispute is to ensure the multiplier is applied correctly and that you're not charged for usage you didn't incur.
Preventing Future Discrepancies
- Log usage programmatically: Capture the
usagefield from every API response and store it. - Monitor regularly: Set up alerts for unexpected spikes in token consumption.
- Stay informed: Follow provider announcements about pricing or tokenization changes.
- Use the same tokenizer: When possible, use the provider's official tokenizer library for local counting.
By following these steps, you can confidently verify your bill and resolve any billing disputes efficiently.