CLAIJun 17

Learning When to Reason for Text-to-SQL via SFT and DPO

arXiv:2607.22622
Originality Incremental advance
AI Analysis

For practitioners deploying Text-to-SQL systems, this work reduces inference cost without sacrificing accuracy on complex queries.

AutoThinkSQL integrates an auto-thinking mechanism into SFT and DPO for Text-to-SQL, enabling dynamic bypass of reasoning for simple queries and deep CoT for complex ones. On Qwen3-Coder-30B-A3B, it achieves consistent gains on Spider and BIRD while reducing output tokens by 24.6% and 18.3%, and latency by 17.1% and 11.5% over CoT-only generation.

Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inference-time overhead. However, a large fraction of real-world queries are simple lookups or aggregations that can be resolved without multi-step deduction, making forced reasoning wasteful. Thus, we propose AutoThinkSQL, a framework that integrates an auto-thinking mechanism into both Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) on Text-to-SQL. Our approach enables the model to dynamically bypass reasoning for simple queries while invoking deep CoT for complex queries. On Qwen3-Coder-30B-A3B, our method achieves consistent gains compared to the best counterpart baseline on both Spider and BIRD benchmarks while simultaneously reducing average output tokens by 24.6% and 18.3%, and average latency by 17.1% and 11.5% compared to CoT-only generation. Further analysis indicates that the model learns to align its reasoning decisions with query difficulty.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes