Using Qwen for Long Documents: Map-Reduce Splitting Through the Aggregator's /chat/completions
Learn how to process book-length documents with Qwen using a map-reduce approach through a single API key on an aggregator platform. Split the text into chunks, summarize each in parallel, then merge the summaries in a final pass.
Why Long Documents Are Tricky
Qwen models, like most LLMs, have a finite context window. A full book, legal contract, or research paper may exceed that limit by many times. Even if the document fits, sending it in one request can be slow and costly.
The map-reduce pattern solves this: split the document into manageable chunks, process each chunk independently (map), then combine the partial results into a final answer (reduce). With the aggregator's /chat/completions endpoint, you can implement this using a single API key for Qwen and other models.
The Map-Reduce Pattern
- Map: Break the document into chunks of a few thousand tokens each. Send each chunk to Qwen with the same summarization prompt. Collect the summaries.
- Reduce: If the combined summaries still exceed the context window, recursively summarize them. Otherwise, send them all in one final request to produce the overall summary.
This pattern is parallelizable: map requests are independent and can be sent concurrently.
Step 1: Chunk the Document
Split by natural boundaries (paragraphs, sections) to preserve coherence. Avoid cutting mid-sentence. A common strategy is to accumulate paragraphs until a target chunk size is reached, with a small overlap (e.g., one paragraph) to maintain context.
Example in Python:
def chunk_text(text, max_chars=4000, overlap=200):
paragraphs = text.split("\n\n")
chunks = []
current = ""
for para in paragraphs:
if len(current) + len(para) > max_chars and current:
chunks.append(current)
# start new chunk with overlap from previous
current = current[-overlap:] + "\n\n" + para
else:
current += ("\n\n" if current else "") + para
if current:
chunks.append(current)
return chunks
Adjust max_chars based on the model's context window and your cost/latency tolerance.
Step 2: Define a Stable Prompt Template
Use the same prompt for every chunk to get consistent summaries. A template like:
Summarize the following section of a larger document in 3-5 sentences. Focus on key facts, arguments, and conclusions. Do not add commentary.
---
{chunk}
Send this to Qwen via /chat/completions:
import requests
def summarize_chunk(chunk, api_key):
response = requests.post(
"https://api.example.com/v1/chat/completions", # replace with aggregator's endpoint
headers={"Authorization": f"Bearer {api_key}"},
json={
"model": "qwen-plus", # or another Qwen variant
"messages": [
{"role": "user", "content": f"Summarize the following section...\n\n{chunk}"}
],
"temperature": 0.3
}
)
return response.json()["choices"][0]["message"]["content"]
Use a low temperature for factual summarization.
Step 3: Map in Parallel
Send all chunks concurrently to reduce total time. Use a thread pool or async requests:
from concurrent.futures import ThreadPoolExecutor
def map_summaries(chunks, api_key):
with ThreadPoolExecutor(max_workers=5) as executor:
futures = [executor.submit(summarize_chunk, chunk, api_key) for chunk in chunks]
return [f.result() for f in futures]
Be mindful of rate limits; start with a small number of workers and adjust.
Step 4: Reduce to a Final Summary
Combine the partial summaries. If they fit in the context window, send them all at once:
def reduce_summaries(summaries, api_key):
combined = "\n\n".join(summaries)
prompt = f"Combine these section summaries into a single coherent summary:\n\n{combined}"
# call /chat/completions similarly
If the combined text is too long, apply the map-reduce recursively: chunk the summaries, summarize each, and repeat until the result fits.
Practical Tips
- Preserve order: Keep chunks in original order so the final summary flows logically.
- Use a consistent model: Stick with the same Qwen variant for all map steps to avoid style shifts.
- Cache results: Store chunk summaries to avoid recomputation if you need to re-run the reduce step.
- Handle errors: Retry failed chunk requests with exponential backoff.
- Monitor tokens: Track usage to manage costs; the aggregator bills at official price × 1.3.
Using the Aggregator's Benefits
With one API key, you can call Qwen and other models (Claude, GPT, DeepSeek, etc.) without separate accounts. Top up with USDC on Base, no KYC. If you contribute keys, you earn credits at official price × 1.1 (or × 1.2 for premium) in USDC.
When to Use This Pattern
Map-reduce is ideal for:
- Books, long reports, or transcripts
- Multi-document summarization
- Any task where the input exceeds the context window
For shorter texts, a single request is simpler. But for book-length documents, map-reduce with Qwen through the aggregator's /chat/completions endpoint provides a scalable, cost-effective solution.