Semantically Similar, Logically Distinct: Diagnosing the Semantic-Answerability Gap in Table RAG
For practitioners building table-based RAG systems, the paper diagnoses a fundamental mismatch between semantic retrieval and answerability, and shows that a simple verification step can largely close the gap.
The paper identifies a 'Semantic-Answerability Gap' in table RAG, where retrievers select semantically similar tables that lack the evidence needed to answer a query. On the new TCR-Bench benchmark, top-5 retrieval drops QA accuracy from 0.755 (oracle) to 0.330, but a lightweight answerability-aware reranking step recovers top-1 target retrieval from 18.2% to 57.4%.
Tables are a critical knowledge source in retrieval-augmented generation (RAG), but a retrieved table may lack sufficient evidence to answer a query, a property we call answerability. While answerability broadly concerns whether a source or collection of sources contains sufficient evidence, retrieval models optimized for semantic relevance do not guarantee it even in the single-source case, creating a fundamental mismatch. To study this, we introduce TCR-Bench, a diagnostic benchmark for Table Content-level Answerability in RAG, built around sibling tables, i.e., tables with highly similar schemas but subtle content differences. On TCR-Bench, the dense retrievers we evaluate persistently exhibit a Semantic-Answerability Gap: they often retrieve the correct sibling group yet struggle to pinpoint the uniquely answerable table within it, dropping QA performance from 0.755 (oracle) to 0.330 (top-5 retrieved). Our analysis suggests this gap is associated with semantic accumulation, schema-level cue dependence, and weak row-column binding. As a diagnostic probe into the source of this gap, we test whether a lightweight two-stage pipeline, Answerability-Aware Reranking (AAR), applying direct query-table answerability judgment, can recover performance: it raises top-1 target retrieval from 18.2% to 57.4%, and this large gain is itself evidence that much of the observed failure reflects a missing answerability verification step, rather than an inherent limitation of model capacity alone.