Extracting Fields from Handwritten Forms and Receipts with a Vision Model
Learn how to extract structured fields from handwritten forms and receipts using a vision model, even when scans are rotated or low-quality. This guide covers image preprocessing, prompt framing for field extraction, and JSON validation.
Introduction
Extracting data from handwritten forms and receipts can be challenging, especially when scans are low-quality or rotated. Vision models can help automate this process, but getting accurate results requires careful preprocessing, prompt design, and validation. This guide walks you through the steps to reliably extract fields using a vision model via an LLM API aggregator.
Why Handwritten Forms and Receipts Are Hard
Handwritten text varies greatly in style, size, and clarity. Scans may be skewed, faded, or have background noise. Receipts often have small fonts, crumpled paper, or poor lighting. These factors make OCR and field extraction error-prone. A vision model, combined with good preprocessing and prompting, can overcome many of these issues.
Image Preprocessing
Before sending an image to a vision model, preprocess it to improve legibility:
- Deskew and rotate: Use image processing libraries (e.g., OpenCV) to detect and correct rotation. Even a small tilt can reduce accuracy.
- Enhance contrast: Convert to grayscale and apply adaptive thresholding or contrast stretching to make text stand out.
- Denoise: Remove background noise with filters like Gaussian blur or median blur.
- Resize: Ensure the image is large enough for the model to read details, but not excessively large to avoid token limits.
- Crop to region of interest: If you only need specific fields, crop the relevant area to reduce distractions.
Prompt Framing for Field Extraction
How you frame the prompt greatly affects the model's output. Follow these best practices:
- Be explicit about the fields: List the exact field names you want extracted (e.g., "invoicenumber", "totalamount", "date").
- Specify the output format: Request a JSON object with a defined structure. For example:
{"field1": "value", "field2": "value"}. - Handle uncertainty: Instruct the model to return
nullfor missing or illegible fields. - Provide context: Mention that the image may be handwritten, rotated, or low-quality, and ask the model to do its best.
- Use system messages: If the API supports system prompts, set a role like "You are an expert at extracting data from scanned documents."
Example prompt:
Extract the following fields from the provided image: invoice_number, date, total_amount, vendor_name. Return a JSON object with these keys. If a field is not present or unreadable, set its value to null. The image may be handwritten or slightly rotated.
Validating the Returned JSON
Vision models can occasionally return malformed JSON or hallucinate values. Always validate and sanitize:
- Parse JSON: Use a strict JSON parser. If parsing fails, attempt to extract JSON from the response (e.g., find the first
{and last}). - Check schema: Ensure all required fields are present and have the correct data types.
- Sanitize values: Trim whitespace, validate date formats, and check numeric ranges.
- Cross-verify: For critical fields (e.g., totals), consider a second pass or manual review.
- Handle errors gracefully: If validation fails, log the response and retry with a modified prompt or preprocessing.
Putting It All Together
Here's a typical workflow:
- Preprocess the image (deskew, enhance, denoise).
- Send the image to a vision model with a well-crafted prompt.
- Receive the JSON response.
- Validate and sanitize the JSON.
- Store or use the extracted fields.
Choosing a Vision Model
Many models support vision inputs. When using an LLM API aggregator, you can access multiple models (e.g., Claude, GPT, DeepSeek) with a single API key. Test different models to see which performs best on your specific documents. Keep in mind that pricing varies, and the aggregator applies a markup (users pay official price × 1.3).
Conclusion
Extracting fields from handwritten forms and receipts is feasible with vision models, but success depends on preprocessing, prompt engineering, and validation. By following these steps, you can build a robust pipeline that handles real-world imperfections.