HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments
To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning.
ProofPaper ↗
Key points
- Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses.
- This creates a critical train-deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time.
- HarnessSQL builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL.
- Across Spider 2.0-SQLite, HarnessSQL dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench.
Sources (1)
- [1]HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database EnvironmentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:36 PM
To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning.
Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
- Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
- Oct 8, 2026Opera: A Verbal Critic Framework for Long-horizon Coding Agents
- Oct 7, 2026Spatial Latent Reasoning for Embodied Reference Understanding
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning
