One Endpoint, Three Tools: Using the Same API Key in a CLI Script, a Notebook and a Chat App
Learn how to configure a CLI script, a Jupyter notebook, and a chat app to use the same API key and base URL from an LLM API aggregator. This guide covers per-client timeout and retry settings, plus strategies to prevent one tool from starving others.
Why Use One Endpoint Across Multiple Tools?
An LLM API aggregator lets you access many models through a single API key and base URL. This simplifies management and billing, but when you use the same key in different clients, you need to handle each tool's behavior to avoid conflicts.
Understanding the Shared Key Model
The aggregator provides one API key that works with any compatible client. You pay the official price × 1.3 for usage, and key contributors get credited official price × 1.1 (premium × 1.2) in USDC. There's no KYC, and you top up with USDC on Base.
Common Pitfalls When Sharing a Key
- Uneven retry storms: A CLI script with aggressive retries can flood the API and delay responses for your notebook.
- Timeout mismatches: A chat app expecting sub-second responses may time out if a notebook runs long batch jobs on the same key.
- Concurrency limits: Some clients open many parallel connections by default, consuming all available slots.
Configuring a CLI Script
Most CLI tools let you set the base URL and key via environment variables or flags.
- Set a reasonable timeout: avoid infinite waits; 30–60 seconds is common for LLM calls.
- Configure retries with exponential backoff: start with 1–2 retries to avoid overwhelming the API.
- Limit concurrency: if the tool supports parallel requests, cap it at a small number (e.g., 2–4).
Example environment setup:
export OPENAI_API_BASE="https://your-aggregator.com/v1"
export OPENAI_API_KEY="sk-..."
Then call your CLI tool, ensuring it respects these settings.
Configuring a Notebook
In Jupyter or similar, you often use HTTP libraries directly or SDKs like openai.
- Set a per-request timeout to avoid hanging cells.
- Implement retry logic with jitter to prevent synchronized retries across notebook runs.
- Consider using asynchronous requests with a semaphore to limit concurrent calls.
Python example with httpx:
import httpx
import asyncio
async def call_model(prompt):
async with httpx.AsyncClient(timeout=30.0) as client:
resp = await client.post(
"https://your-aggregator.com/v1/chat/completions",
headers={"Authorization": "Bearer sk-..."},
json={"model": "claude-3", "messages": [{"role": "user", "content": prompt}]},
)
resp.raise_for_status()
return resp.json()
Add retries with tenacity or similar.
Configuring a Chat App
GUI-based chat apps (e.g., Open WebUI, LM Studio) typically have settings for base URL and API key.
- Set a shorter timeout for interactive use (e.g., 15–30 seconds).
- Reduce retry attempts to avoid UI freezes.
- Disable streaming if the app struggles with long responses.
Avoiding Cross-Tool Starvation
When all tools share one key, they compete for the same rate limits. To keep things smooth:
- Stagger schedules: Don't run heavy notebook batches while actively using the chat app.
- Use separate keys if possible: Some aggregators allow multiple keys per account; assign one per tool to isolate rate limits.
- Monitor usage: Keep an eye on your aggregator dashboard to spot which tool consumes most.
- Implement client-side rate limiting: In each tool, add a small delay between requests (e.g., 100 ms) to spread load.
Conclusion
Sharing one API key across a CLI, notebook, and chat app is convenient but requires mindful configuration. Set appropriate timeouts and retries per client, limit concurrency, and consider separate keys if the aggregator supports it. This way, you get the benefits of a unified endpoint without sacrificing responsiveness.