AICLJun 18

Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

arXiv:2606.198087.8
Predicted impact top 78% in AI · last 90 daysOriginality Synthesis-oriented
AI Analysis

For practitioners deploying LLMs with test-time reasoning, this provides a deployment rule: tune initial budget first, then use selective recovery when explicit checks or risk control matter, but the gains are incremental and often surpassed by simply increasing initial compute.

The paper studies selective verification for budget-aware reasoning, introducing SeVRA, a controller that decides whether to preserve a frozen solver's initial answer or invoke active verification. On MATH500, it reaches 76.3% accuracy vs 75.5% for always verifying, reducing post-generation tokens by 26.8% and harmful flips from 2.2% to 1.0%, though a longer initial solve achieves 76.0% with 28% fewer total tokens.

Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-correct answers, or introduce harmful answer changes. We study this as a deployment allocation problem rather than a new-verifier problem. We introduce \sevra, Selective Verification for Reasoning Allocation, a serving-layer controller that decides whether to preserve a frozen solver's initial answer or invoke active verification. Using a frozen Qwen3-4B solver, we log intervention outcomes and train recoverability-aware gates from serving-visible attempt state. On \mathfive, selective verification reaches 76.3\% accuracy, compared with 75.5\% for always verifying, while reducing post-generation tokens by 26.8\% and harmful flips from 2.2\% to 1.0\%. However, an 8,192-token initial solve reaches 76.0\% accuracy with 28\% fewer total model tokens, showing that selective recovery is useful but not the best tested cost frontier. In frozen transfer to \gsm, the selective policy verifies only 3.0\% of examples, improves accuracy from 93.4\% to 94.5\%, and reduces verification tokens by 91.2\% relative to always verifying; again, a longer initial solve matches its accuracy with fewer realized tokens. On CommonsenseQA, always-on verification hurts, while Self-Consistency@5 improves accuracy at about five times the realized token cost. The resulting deployment rule is: tune the initial budget first, then use selective recovery when explicit checks, bounded retries, auditability, or regression-risk control matter.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes