ResearchResearch paperLarge Language Models · Speech & Audio · Efficiency & Inference1 source · Oct 8, 2026

SP-DocReader: Difference-Aware Self-Play for Precise Document OCR

We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.

Key points

  • Accurate page transcription remains difficult for vision language models under limited input and training budgets.
  • Reading Discrepancy Masking aligns reference and generated model tokens through a longest common subsequence, then scores unmatched positions with their full conditioning prefixes.
  • Focused Fidelity Loss adds direct negative log-likelihood supervision at unmatched ground-truth positions.
  • On Qwen3-VL-4B, it reduces character error rate by approximately 54 percent and improves DocVQA Average Normalized Levenshtein Similarity (ANLS) by 3.7 points.

Sources (1)

  • [1]SP-DocReader: Difference-Aware Self-Play for Precise Document OCR
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:12 AM
    We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.
    Accurate page transcription remains difficult for vision language models under limited input and training budgets.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
  2. Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
  3. Oct 8, 2026Opera: A Verbal Critic Framework for Long-horizon Coding Agents
  4. Oct 7, 2026Spatial Latent Reasoning for Embodied Reference Understanding
  5. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  6. Oct 6, 2026The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning

Related