Measuring Time-to-First-Token vs Total Duration on Aggregated Models
Learn how to measure time-to-first-token (TTFT) and total duration separately for models accessed through an aggregated API. This guide explains why both metrics matter for perceived performance, how to log them correctly in streaming and non-streaming modes, and how to use them to choose models and optimize user experience.
When you call a language model through an aggregated API, the overall request duration includes several phases: network latency to the provider, queueing, prompt processing, and token generation. Time-to-first-token (TTFT) measures how long it takes before the first token of the response arrives, while total duration covers the entire request from start to finish.
Why Measure Both
Users perceive streaming responses as fast when the first token appears quickly, even if the full answer takes longer. By logging TTFT and total duration separately, you can:
- Track perceived responsiveness: TTFT is a strong indicator of how snappy the interaction feels.
- Diagnose bottlenecks: A high TTFT suggests prompt processing or queueing delays; a high total duration with low TTFT indicates slow token generation.
- Compare models fairly: Different models may excel at different phases. For example, one model might have low TTFT but slower throughput, while another is the opposite.
Understanding the Phases
A typical API call goes through these stages:
- Client sends request over the network to the aggregator.
- Aggregator routes the request to the selected model provider.
- Provider processes the prompt (prefill phase).
- Provider generates tokens one by one (decode phase).
- Tokens stream back to the client (or arrive in a single batch if non-streaming).
TTFT is the time from sending the request to receiving the first token. Total duration is the time from sending the request to receiving the last token (or the complete response in non-streaming mode).
How to Log TTFT and Total Duration
The implementation depends on whether you use streaming or non-streaming calls.
Streaming Mode
In streaming, you receive tokens incrementally. Log the timestamp when you send the request, when the first chunk arrives, and when the last chunk arrives.
- Start time: Record immediately before sending the request.
- First token time: Record when the first token (or first content chunk) is received.
- End time: Record when the stream closes or the final chunk arrives.
TTFT = first token time − start time. Total duration = end time − start time.
Non-Streaming Mode
In non-streaming, you only get the full response at once. You cannot measure TTFT directly because the first token is part of the complete response. However, you can still log total duration. To approximate TTFT, you would need to switch to streaming or rely on provider-specific headers if available.
Best Practices
- Log per model: Store TTFT and total duration alongside the model name, so you can compare performance across models.
- Include metadata: Record input token count, output token count, and whether streaming was used. This helps interpret the numbers.
- Use percentiles: Averages can hide outliers. Track median and 95th percentile for both metrics.
- Separate network latency: If possible, measure the time to establish a connection and the time to first byte to isolate provider processing time.
Using the Metrics
With TTFT and total duration logged, you can make informed decisions:
- Model selection: If your application requires fast initial feedback (e.g., chatbots), prioritize models with low TTFT. If you need high throughput for long outputs, consider total duration.
- Performance tuning: If TTFT is high, try reducing prompt length or check if the provider is experiencing high load. If total duration is high but TTFT is low, the model may be slow at generating tokens; consider a different model for long-form content.
- User experience: In streaming UIs, show a typing indicator until the first token arrives. Once tokens start flowing, the user perceives progress even if the total time is longer.
Example Logging Approach
Here’s a conceptual example in pseudocode:
start = now()
response = call_model(prompt, stream=True)
first_token_time = None
for chunk in response:
if first_token_time is None:
first_token_time = now()
process_chunk(chunk)
end = now()
ttft = first_token_time - start
total_duration = end - start
log(model=model_name, ttft=ttft, total_duration=total_duration, ...)
Conclusion
Measuring TTFT and total duration separately gives you a clearer picture of model performance and user-perceived speed. By logging these metrics per model, you can optimize your application to feel fast even when total generation time is longer. Remember that streaming is essential for a responsive experience, and TTFT is your key indicator for that.