ResearchResearch paperLarge Language Models1 source · Oct 8, 2026

HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments

To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning.

Key points

  • Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses.
  • This creates a critical train-deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time.
  • HarnessSQL builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL.
  • Across Spider 2.0-SQLite, HarnessSQL dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench.

Sources (1)

  • [1]HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:36 PM
    To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning.
    Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
  2. Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
  3. Oct 8, 2026Opera: A Verbal Critic Framework for Long-horizon Coding Agents
  4. Oct 7, 2026Spatial Latent Reasoning for Embodied Reference Understanding
  5. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  6. Oct 6, 2026The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning

Related