parallel-tool-calls-single-turn · EN · 2026-10-08

Parallel Tool Calls in One Turn: Fan-Out Tools and Reassemble Results

Learn how to enable parallel tool calls in a single assistant turn, execute them concurrently, and reassemble results for a follow-up call. This guide covers the mechanics, benefits, and practical implementation steps, with considerations for LLM API aggregators.

Why Parallel Tool Calls Matter

When an assistant needs information from multiple sources—such as checking weather in two cities or fetching data from several APIs—executing tools sequentially can slow down responses and increase latency. Parallel tool calls allow the model to request multiple tools in one turn, which you can execute concurrently and then return all results for a single follow-up. This reduces round-trips and improves efficiency.

How a Single Turn Requests Multiple Tools

Modern LLM APIs support tool calling where the model can emit multiple tool call requests in one response. Each tool call includes:

  • A unique id to correlate the call with its result.
  • A type (usually "function").
  • A function object containing the name and arguments.

For example, a user asks: "What's the weather in Paris and London?" The model might respond with two tool calls in a single message: one for Paris, one for London.

Executing Tools Concurrently

Since tool calls are independent, you can execute them in parallel. The exact method depends on your environment:

  • Async/await: Use asyncio.gather or similar to run multiple async functions concurrently.
  • Threads: Use a thread pool to execute blocking calls in parallel.
  • Promises: In JavaScript, use Promise.all.

The key is to avoid sequential execution unless there are dependencies between tools.

Reassembling Results into a Follow-Up Call

After executing all tools, you must send the results back to the model in a single follow-up request. The conversation history should include:

  1. The original user message.
  2. The assistant message containing the tool calls.
  3. A new message for each tool result, with role "tool", the corresponding toolcallid, and the output.

All tool result messages must be sent together in one API call. The model then uses these results to generate a final answer.

Handling Errors and Timeouts

If a tool fails, include the error message in the tool result. The model can then decide whether to retry or proceed. Set reasonable timeouts for each tool to prevent one slow call from blocking others. If a timeout occurs, return an error result so the model can handle it gracefully.

Best Practices

  • Idempotency: Ensure tools are safe to retry if needed.
  • Rate limits: Be mindful of API rate limits when fanning out many calls.
  • Context length: Parallel tool calls increase token usage; monitor context limits.
  • User experience: Inform users if tools take time; consider streaming partial results.

Using an LLM API Aggregator

If you're using an LLM API aggregator that provides a unified interface to multiple models (such as Claude, GPT, DeepSeek, Qwen, GLM, Kimi), the parallel tool calling workflow remains the same. The aggregator handles routing to the appropriate model and provides a consistent API. With a single API key, you can access various models and pay in USDC on Base without KYC. The cost is official price × 1.3 for users, and contributors are credited official price × 1.1 (premium × 1.2) in USDC.

Example Flow

  1. User asks a multi-part question.
  2. Model returns multiple tool calls.
  3. Your code executes them concurrently.
  4. You send all tool results in one follow-up message.
  5. Model returns the final answer.

This pattern is efficient and scales well for tasks requiring multiple data fetches or actions.