Parallel Tool Calls in One Turn: Fan-Out Tools and Reassemble Results
Learn how to enable parallel tool calls in a single assistant turn, execute them concurrently, and reassemble results for a follow-up call. This guide covers the mechanics, benefits, and practical implementation steps, with considerations for LLM API aggregators.
Why Parallel Tool Calls Matter
When an assistant needs information from multiple sources—such as checking weather in two cities or fetching data from several APIs—executing tools sequentially can slow down responses and increase latency. Parallel tool calls allow the model to request multiple tools in one turn, which you can execute concurrently and then return all results for a single follow-up. This reduces round-trips and improves efficiency.
How a Single Turn Requests Multiple Tools
Modern LLM APIs support tool calling where the model can emit multiple tool call requests in one response. Each tool call includes:
- A unique
idto correlate the call with its result. - A
type(usually "function"). - A
functionobject containing thenameandarguments.
For example, a user asks: "What's the weather in Paris and London?" The model might respond with two tool calls in a single message: one for Paris, one for London.
Executing Tools Concurrently
Since tool calls are independent, you can execute them in parallel. The exact method depends on your environment:
- Async/await: Use
asyncio.gatheror similar to run multiple async functions concurrently. - Threads: Use a thread pool to execute blocking calls in parallel.
- Promises: In JavaScript, use
Promise.all.
The key is to avoid sequential execution unless there are dependencies between tools.
Reassembling Results into a Follow-Up Call
After executing all tools, you must send the results back to the model in a single follow-up request. The conversation history should include:
- The original user message.
- The assistant message containing the tool calls.
- A new message for each tool result, with role "tool", the corresponding
toolcallid, and the output.
All tool result messages must be sent together in one API call. The model then uses these results to generate a final answer.
Handling Errors and Timeouts
If a tool fails, include the error message in the tool result. The model can then decide whether to retry or proceed. Set reasonable timeouts for each tool to prevent one slow call from blocking others. If a timeout occurs, return an error result so the model can handle it gracefully.
Best Practices
- Idempotency: Ensure tools are safe to retry if needed.
- Rate limits: Be mindful of API rate limits when fanning out many calls.
- Context length: Parallel tool calls increase token usage; monitor context limits.
- User experience: Inform users if tools take time; consider streaming partial results.
Using an LLM API Aggregator
If you're using an LLM API aggregator that provides a unified interface to multiple models (such as Claude, GPT, DeepSeek, Qwen, GLM, Kimi), the parallel tool calling workflow remains the same. The aggregator handles routing to the appropriate model and provides a consistent API. With a single API key, you can access various models and pay in USDC on Base without KYC. The cost is official price × 1.3 for users, and contributors are credited official price × 1.1 (premium × 1.2) in USDC.
Example Flow
- User asks a multi-part question.
- Model returns multiple tool calls.
- Your code executes them concurrently.
- You send all tool results in one follow-up message.
- Model returns the final answer.
This pattern is efficient and scales well for tasks requiring multiple data fetches or actions.