AION
Research paperLarge Language Models1 source · Oct 8, 2026

Structure Tax: How Structured Output affects LLMs Performance

Deploying large language models in production often requires constraining outputs to structured formats such as JSON or XML, and prior work treats the resulting accuracy loss as an inherent structure tax'.

Key points

  • We re-examine this claim by evaluating a battery of models, datasets and schemas, measuring task accuracy, confidence calibration, and hidden-state geometry.
  • The tax turns out to depend on schema design rather than on structure per se: reasoning-first field ordering matches or exceeds free-form accuracy, while answer-first ordering causes steep drops, particularly in smaller models.
  • Format sensitivity scales inversely with a task's own structural constraints, and schemas that preserve reasoning order also improve calibration with CKA showing greater separability between correct and incorrect representations in middle transformer layers.
  • Our findings indicate that properly designed structured formats can match or exceed free-form performance, reframing the critical question from whether to structure' to how to structure' for optimal reasoning preservation.

Sources (1)

  • [1]Structure Tax: How Structured Output affects LLMs Performance
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:39 PM
    Deploying large language models in production often requires constraining outputs to structured formats such as JSON or XML, and prior work treats the resulting accuracy loss as an inherent `structure tax'.
    We re-examine this claim by evaluating a battery of models, datasets and schemas, measuring task accuracy, confidence calibration, and hidden-state geometry.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
  2. Oct 8, 2026One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
  3. Oct 8, 2026Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
  4. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  5. Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
  6. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0

Related