AION
Research paperLarge Language Models · Computer Vision · Multimodal Models1 source · Oct 8, 2026

From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction

We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections.

Key points

  • While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing.
  • This is especially acute in heritage digitization, where documents include historical handwriting, domain-specific terminology (e.g., jewelry, prehistory, architecture), and non-standard layouts requiring high-dimensional structured extraction.
  • We evaluate models under a constraint-aware protocol across zero-shot, few-shot, and fine-tuning settings, measuring extraction fidelity and structured-output quality using Character Error Rate (CER), Approximate Normalized Levenshtein Similarity (ANLS), and mean Average Precision F1 (mAP-F1).
  • Overall, we show that carefully adapted VLMs with up to 7B parameters can provide a sustainable, private, high-performing alternative to manual transcription or commercial black-box systems, and we offer actionable guidance for heritage institutions seeking institution-controlled OCR-to-JSON extraction.

Sources (1)

  • [1]From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 12:21 PM
    We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections.
    While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
  2. Oct 8, 2026SuperNav: An Agentic Navigation System for Any Task in Any Scene
  3. Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
  4. Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
  5. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  6. Oct 7, 2026Q-Learning with Scalar Adjoint Matching

Related