Semantically Similar, Yet Not Answerable: Diagnosing the Semantic-Answerability Gap in Table RAG

arXiv:2607.17742 · cs.AI · Submitted 2026-07-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Semantically Similar, Yet Not Answerable".

Jane: Tables are critical knowledge sources in retrieval-augmented generation (RAG), but retrieved tables often lack sufficient evidence to answer a query, creating a fundamental mismatch between semantic relevance and answerability.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let's look at the title of this paper: "Semantically Similar, Yet Not Answerable: Diagnosing the Semantic-Answerability Gap in Table RAG." It really nails the core problem we're seeing in table retrieval where relevance doesn't translate into actual answerability.

Jane: And the authors are Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang, Lihua Yu, Zujie Ren, Gang Chen, and Junbo Zhao from Zhejiang University and Bank of Hangzhou Co., Ltd. They seem to have put together a very focused diagnostic tool for this issue.

Lu: The authors are tackling the problem systematically by introducing TCR-Bench to isolate answerability from the general semantic similarity that dense retrievers usually optimize for. That diagnostic approach is really smart because it lets you control exactly what you're testing.

Meng: Isolating the variables is crucial, I think. If they can prove that the issue isn't just a random failure but something tied to specific structural or binding issues, then we know where to focus our engineering efforts for real improvements in accuracy.

Lalam: When they study this gap, it helps us understand how much trust we can place in AI when it’s looking at structured data; if the system can't verify the evidence, that confidence level drops significantly.

The paper's summary: Tom: The paper summarizes that dense retrievers often capture topical signal but fail to distinguish between tables with highly similar schemas and subtle content differences, which is what they term the Semantic-Answerability Gap. They show this happens even when the retrieval model gets the correct general topic.

Jane: So, basically, if you ask a question about Table A and it retrieves Table B because they are semantically close, but Table B doesn't actually contain the specific answer you need, that’s the gap they are describing. It’s a mismatch between what looks similar and what is actually verifiable.

Lu: They pinpoint three main drivers behind this gap: semantic accumulation, schema-level cue dependence where models over-rely on high-level structure, and weak row-column binding in how embeddings encode relational structure. That gives us a very concrete list of technical culprits to investigate.

Meng: Schema-level cue dependence is interesting because it suggests the model is getting distracted by the table layout rather than focusing on the specific data within that layout for answering the query. That points toward a need for better structural awareness in our embedding models.

Lalam: If schema cues are causing this, it means we need to make sure our AI doesn't get confused by how a table is formatted or organized, but rather focus on the specific content relationships that matter for answering questions.

The paper's improvements: Tom: The key improvement they propose is Answerability-Aware Reranking, or AAR. Instead of just relying on the initial retrieval score, this two-stage pipeline introduces explicit answerability judgment during reranking to pinpoint the right source more accurately.

Jane: That’s a big step because it moves beyond just finding candidates to actively judging which candidate actually has the evidence for your query before passing it along for final generation. They show that this helps raise top-one retrieval from eighteen point two percent up to fifty-seven point four percent.

Lu: The results of AAR are quite telling; the best DS@one achieved is zero point seven zero four, which they suggest means that this explicit interaction-based answerability modeling substantially improves exact target identification compared to the embedding-only baseline. That’s a measurable gain in precision.

Meng: I see how that works from an engineering viewpoint; adding a secondary check based on answerability judgment adds overhead, but if it significantly boosts the success rate of finding the correct table, that cost seems justifiable for improving downstream QA performance.

Lalam: For our system culture, this means we can build models that are more reliable because they are explicitly checking for evidence before committing to an answer; it builds a layer of verifiable certainty into our AI's thinking process.

Conclusion: Tom: So, to wrap up the findings from "Semantically Similar, Yet Not Answerable: Diagnosing the Semantic-Answerability Gap in Table RAG," the paper concludes that dense retrievers achieve strong coarse semantic relevance but exhibit near-chance fine-grained answerability on their own.

Jane: That means we need to move beyond just matching similarity; we need a mechanism that verifies if the retrieved source actually contains enough evidence for the specific query being asked. They stress that query-aware reranking substantially alleviates this issue by restoring answerability verification that single-vector retrieval misses.

Lu: The implication is clear: future work should focus on embedding content and structure awareness earlier in the pipeline to balance efficiency with the need for precise answerability. They also highlighted that row-level retrieval can exceed zero point eight seven R@one but table-level selection remains below zero point two six, showing where the limitation truly lies.

Meng: From an engineering perspective, this suggests we should explore multi-vector retrieval approaches where rows are embedded independently to better capture local key-value associations before trying to aggregate them for a final table decision.

Lalam: This work really reinforces that precise answerability at top ranks is more important than broad semantic recall when dealing with tables; semantically related distractors can seriously interfere with the downstream reasoning process if they aren't properly filtered out.

Zhejiang University

cs.AI

Submitted: 2026-07-20

Updated: 2026-09-28

Code: https://github.com/minger-hsxz/TCR-Bench-Open

Importance score: 83/100

The gist: Tables are critical knowledge sources in retrieval-augmented generation (RAG), but retrieved tables often lack sufficient evidence to answer a query, creating a fundamental mismatch between semantic

Key concepts

Semantic-Answerability Gap
This is a problem where retrieval systems capture general topic similarity but cannot distinguish the single correct table needed to answer a specific question. Retrievers often pull in many similar tables, making it hard for the system to find the exact evidence required.
TCR-Bench
A diagnostic benchmark designed around 'sibling tables'—tables with very similar structures but slightly different content. This setup isolates whether retrieval fails due to coarse semantic similarity or lack of precise answerability.
Answerability-Aware Reranking (AAR)
A two-stage mitigation strategy where an explicit judgment is added during the reranking phase. It forces the model to evaluate how well a retrieved table actually answers the specific query, leading to much higher precision in selecting the correct target table.

Terminology

Summary

Tables are critical knowledge sources in retrieval-augmented generation (RAG), but retrieved tables often lack sufficient evidence to answer a query, creating a fundamental mismatch between semantic relevance and answerability.

The Semantic-Answerability Gap is identified as a systemic failure mode where dense retrievers capture coarse semantic similarity but fail to distinguish the uniquely answerable table within semantically similar groups. This gap leads to a significant drop in QA performance, as retrieval models often struggle to pinpoint the exact evidence-bearing source despite retrieving the correct topical neighborhood.

How it works

The core of the study is TCR-Bench, a diagnostic benchmark built around sibling tables—tables with highly similar schemas but subtle content differences. This setup isolates answerability from coarse semantic relevance and enables controlled analysis of retrieval behavior. The benchmark constructs queries associated with one Target Table and multiple Sibling Distractor Tables derived from the same source table, ensuring exactly one valid Target Table per query.

The paper identifies three mechanisms driving this gap:

  1. Semantic accumulation: Retrievers capture topical signal rather than verifying precise evidence required to answer a query.

  2. Schema-level cue dependence: Models over-rely on high-level schema signals, limiting discrimination among structurally similar tables.

  3. Weak row-column binding: Embeddings weakly encode relational structure, failing to robustly preserve row-column bindings required for answerability, as evidenced by the limited penalty incurred when column shuffling is applied.

Key Findings and Diagnostics

The study systematically probes surface variations to rule out superficial artifacts:

** Effect of Table Serialization Format**

Retrieval performance remains relatively stable across formats (Markdown, CSV, HTML), indicating that the gap is not caused by surface-level formatting or serialization artifacts. Consistency metrics show high Hit Consistency and Partial Consistency across format pairs.

** Effect of Query Paraphrasing**

Query paraphrasing changes ranking but preserves highly overlapping candidate sets, suggesting retrievers capture general query intent but lack the fine-grained discrimination required for answerability verification.

Mitigation Strategy: Answerability-Aware Reranking (AAR)

To diagnose and mitigate the gap, the authors introduce Answerability-Aware Reranking (AAR), a lightweight two-stage pipeline that reintroduces explicit answerability judgment. This involves applying direct query-table answerability judgment during reranking. The results show substantial improvement: AAR raises top-1 retrieval from 18.2% to 57.4%, and the best DS@1 achieved is 0.704, suggesting that explicit interaction-based answerability modeling substantially improves exact target identification.

Implications for Retrieval Objectives

The findings suggest that the primary limitation lies in the lack of interaction-based answerability assessment during initial retrieval. The study concludes that the dense retrievers we evaluate achieve strong coarse semantic relevance but exhibit near-chance fine-grained answerability, and Query-aware reranking substantially alleviates this issue. This points toward answerability-aware retrieval as a promising direction worth further investigation, beyond coarse semantic matching.

Conclusion

The Semantic-Answerability Gap is a fundamental bottleneck in Table RAG, stemming from the mismatch between contrastive alignment objectives and fine-grained content verification. The work formalizes this gap and demonstrates that explicit query-table interaction can restore answerability precision. Future systems should embed content- and structure-aware representations earlier to balance efficiency with precise answerability.

The TCR-Bench is publicly released, available at: https://github.com/minger-hsxz/TCR-Bench-Open. The benchmark is open-sourced under the CC BY-SA 4.0 license. It includes table data files, complete query files, and the full implementation of the RAG pipeline used in this paper.

The results show that Answerability precision at top ranks is more important than broad semantic recall; semantically related distractors interfere with downstream reasoning. Furthermore, Row-level retrieval exceeds 0.87 R@1, while table-level Target Table selection remains below 0.26, indicating that local key-value associations are captured but do not scale to reliable table-level answerability discrimination. Finally, Explicit query-table interaction restores answerability verification that is difficult to capture through single-vector semantic retrieval alone.

Table 28 summarizes the major experimental observations regarding answerability-aware table retrieval and the SemanticAnswerability Gap. The best embedding retriever achieves only 18.2% Top-1 Target Table retrieval despite consistently retrieving tables from the correct sibling group, and AAR reranking substantially improves both R@1 and DS@1 compared to the embedding-only Qwen3-8B baseline.

Improvements for AI systems

Based on the scientific paper, here are specific improvements for AI systems and what those improved systems can achieve:


  1. Improve Retrieval Systems by Incorporating Explicit Answerability Verification (Answerability-Aware Reranking - AAR).

  2. Implement a lightweight two-stage pipeline where a primary dense retriever is followed by an answerability-aware reranker (like AAR-CE or AAR-Judge). This reranker should explicitly score candidates based on whether they contain sufficient evidence to answer the query, rather than solely on semantic similarity.

  3. Use this improved retrieval system to achieve a statistically significant increase in Top-1 Target Table retrieval (e.g., raising it from 18.2% to 57.4%) and substantially improving downstream QA performance (F1 score).

  4. Develop Retrieval Architectures that Prioritize Row-Column Binding over General Semantic Similarity:

  5. Design or fine-tune embeddings/retrievers that are explicitly trained or designed to preserve relational structure, specifically focusing on:

  6. Implementing robust mechanisms against Weak Row-Column Bindings, such as incorporating explicit positional signals (like row indices) and testing the effect of column shuffling to see how much performance degrades.

  7. Utilize Multi-vector Retrieval approaches where each table row is independently embedded, allowing for more precise aggregation strategies (e.g., using max aggregation) that better capture local key-value associations than single-vector retrieval.

  8. Enhance Table Understanding by Incorporating Schema and Structural Cues into Embeddings:

  9. Improve embedding models to be less prone to Schema-Level Cue Dependence by ensuring they do not over-rely on high-level schema signals when discriminating between structurally similar tables with different factual values.

  10. Systemic Improvement for LongtableBench QA:

  11. When performing Question Answering on retrieved tables, the downstream LLM must be equipped with a highly structured prompt (like Table 10) that enforces strict formatting constraints, including numerical normalization (Roman to Arabic), handling of derived columns (e.g., reversing mathematical operations), and date standardization. This ensures the QA model can reliably parse and synthesize complex, structured data from retrieved sources.

  12. Investigate and Mitigate Positional Bias in Table Embeddings:

  13. Develop methods to neutralize the observed Header Preference bias where models favor information appearing near the beginning of a serialized table. This could involve fine-tuning or architectural modifications to give equal weight to different positions within the table structure, preventing positional heuristics from overriding genuine relational reasoning.

Sources

Related papers