Forcing Qwen to Return Clean JSON: Schema Prompts and Retry Loops on an OpenAI-Compatible Endpoint
Learn how to reliably extract structured JSON from Qwen models using schema prompts and bounded retry loops on an OpenAI-compatible endpoint. This guide covers prompt design, validation, and error handling to ensure clean, parseable output for your pipelines.
Why Structured Output Matters
When integrating LLMs into production pipelines, you often need responses in a machine-readable format like JSON. Free-form text is error-prone to parse. Qwen models can produce JSON, but without proper guidance, they may include extra text, use incorrect syntax, or deviate from your schema. This guide shows how to enforce clean JSON output using schema prompts and bounded retry loops.
Schema Prompting
A schema prompt explicitly describes the desired JSON structure. Include it in the system message or as part of the user prompt. Be precise about field names, types, and allowed values.
Example Schema Prompt
You must respond with a JSON object that matches this schema:
{
"name": string,
"age": integer,
"email": string
}
Do not include any other text.
Key practices:
- Use JSON Schema or a simplified type notation.
- Specify required fields and their types.
- For enums, list allowed values.
- Instruct the model to output only the JSON object, without markdown fences or explanations.
Handling Nested Objects
For complex schemas, define nested structures clearly. Use indentation or reference definitions. For example:
{
"user": {
"id": integer,
"profile": {
"bio": string,
"links": array of strings
}
}
}
Validation and Parsing
After receiving the model's response, validate it against your schema. Use a JSON parser and a schema validator like Ajv (JavaScript) or Pydantic (Python).
Common Issues
- Markdown code fences: The model may wrap JSON in ``
json ...``. Strip these before parsing. - Trailing commas: Invalid JSON that parsers reject.
- Missing fields: The model may omit required keys.
- Type mismatches: e.g., a number as a string.
Validation Steps
- Extract the JSON substring: find the first
{and the last}. - Parse with a strict JSON parser.
- Validate against your schema.
- If validation fails, trigger a retry.
Bounded Retry Loop
A retry loop attempts to correct the model's output by re-prompting with the error. Bound the number of retries to avoid infinite loops.
Retry Strategy
- Max retries: Set a limit (e.g., 3).
- Error feedback: Include the validation error in the retry prompt.
- Temperature: Lower temperature (e.g., 0.2) for more deterministic output.
Example Retry Prompt
Your previous response was invalid. Error: Missing required field 'email'.
Please provide a valid JSON object matching the schema.
Pseudocode
def get_structured_output(prompt, schema, max_retries=3):
for attempt in range(max_retries):
response = call_model(prompt)
try:
data = parse_and_validate(response, schema)
return data
except ValidationError as e:
prompt = f"{prompt}\n\nPrevious error: {e}\nProvide valid JSON."
raise Exception("Failed to get valid JSON after retries")
Best Practices
- Be explicit: Clearly state the output format in the prompt.
- Use few-shot examples: Show a correct output for similar input.
- Set stop sequences: If the API supports it, stop at the closing brace.
- Monitor failures: Log validation errors to improve prompts.
- Test edge cases: Ensure your schema handles empty or unexpected inputs.
Conclusion
By combining clear schema prompts with validation and bounded retries, you can reliably obtain clean JSON from Qwen models. This approach works on any OpenAI-compatible endpoint, making it easy to integrate with your existing tooling.