SP-DocReader: Difference-Aware Self-Play for Precise Document OCR
We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.
ProofPaper ↗
Key points
- Accurate page transcription remains difficult for vision language models under limited input and training budgets.
- Reading Discrepancy Masking aligns reference and generated model tokens through a longest common subsequence, then scores unmatched positions with their full conditioning prefixes.
- Focused Fidelity Loss adds direct negative log-likelihood supervision at unmatched ground-truth positions.
- On Qwen3-VL-4B, it reduces character error rate by approximately 54 percent and improves DocVQA Average Normalized Levenshtein Similarity (ANLS) by 3.7 points.
Sources (1)
- [1]SP-DocReader: Difference-Aware Self-Play for Precise Document OCRarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:12 AM
We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.
Accurate page transcription remains difficult for vision language models under limited input and training budgets.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
- Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
- Oct 8, 2026Opera: A Verbal Critic Framework for Long-horizon Coding Agents
- Oct 7, 2026Spatial Latent Reasoning for Embodied Reference Understanding
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning