Learning When to Reason for Text-to-SQL via SFT and DPO
For practitioners deploying Text-to-SQL systems, this work reduces inference cost without sacrificing accuracy on complex queries.
AutoThinkSQL integrates an auto-thinking mechanism into SFT and DPO for Text-to-SQL, enabling dynamic bypass of reasoning for simple queries and deep CoT for complex ones. On Qwen3-Coder-30B-A3B, it achieves consistent gains on Spider and BIRD while reducing output tokens by 24.6% and 18.3%, and latency by 17.1% and 11.5% over CoT-only generation.
Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inference-time overhead. However, a large fraction of real-world queries are simple lookups or aggregations that can be resolved without multi-step deduction, making forced reasoning wasteful. Thus, we propose AutoThinkSQL, a framework that integrates an auto-thinking mechanism into both Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) on Text-to-SQL. Our approach enables the model to dynamically bypass reasoning for simple queries while invoking deep CoT for complex queries. On Qwen3-Coder-30B-A3B, our method achieves consistent gains compared to the best counterpart baseline on both Spider and BIRD benchmarks while simultaneously reducing average output tokens by 24.6% and 18.3%, and average latency by 17.1% and 11.5% compared to CoT-only generation. Further analysis indicates that the model learns to align its reasoning decisions with query difficulty.