EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL

arXiv:2605.29670 · cs.CL, cs.AI · Submitted 2026-05-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL".

Tom: Schema linking is reframed as uncertainty-aware schema-need inference over multiple plausible SQL paths, where systems distinguish required schema items from path-dependent uncertain ones and acquire evidence only where needed.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Welcome back to the show everyone! We've got a fascinating paper today on arXiv titled "EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL." Jane, you mentioned this paper tackles a really tricky part of making AI systems understand complex databases. Can you give us the quick rundown of what the core idea is?

Jane: Absolutely, Tom. The main point EviLink makes is that instead of just picking one schema based on one SQL path, it looks at multiple possible SQL paths to answer a question and treats schema linking as an uncertainty-aware inference process across all those possibilities. It claims this approach helps the system correctly separate what schema items are absolutely required from what might just be uncertain depending on which path is taken.

Lu: That sounds incredibly fertile ground for new research, Jane. Thinking about the implications, if a system can robustly handle multiple plausible SQL paths simultaneously, it opens up possibilities where current models fail because they're stuck on a single decision and potentially lose necessary information for other correct solutions.

Meng: From an engineering standpoint, that sounds computationally intensive, Lu. How does this uncertainty awareness translate into something practical when you're building these large-scale Text-to-SQL systems? We need to know how much overhead this adds to the actual inference time.

Lalam: I see a huge cultural shift here, Meng. If we can build systems that don't just follow one rigid script but explore multiple valid reasoning routes, it suggests an AI culture focused on robustness and acknowledging ambiguity rather than just seeking the fastest single answer.

Tom: That’s a big shift, Lalam. So, EviLink essentially moves away from that old idea of deterministic selection around one SQL path, which is what many current methods rely on. Jane, can you elaborate on why this reframing matters for Text-to-SQL?

Jane: Well, the paper points out that existing methods make assumptions about a single SQL path and deterministic decisions when multiple valid implementations exist for a single question. EviLink addresses this by acknowledging that different paths need different schema elements, so linking around just one path can cause irreversible information loss by discarding parts needed by other solutions.

Lu: It’s about recognizing the complexity inherent in real-world data interactions, Jane. Think about how a complex question might have two fundamentally different ways to query the database, and if we only ground our schema on one path, we might miss crucial information required by the other path.

Meng: So if we frame it as inference over multiple paths, how does EviLink actually decide what to keep versus what to discard? Is there a clear mechanism for that decision-making process?

Jane: Yes, the method instantiates this by performing multi-hypothesis schema grounding where the model generates a set of structured hypotheses. It then uses cross-voting to create three sets: a required set, an uncertain set, and a candidate set.

Paper summary: Lalam: That sounds like the system is first trying to map out the landscape before committing to any single path, which is a very thoughtful way for an AI to approach ambiguity. It's building a map of possibilities before choosing a route.

Tom: Building a map of possibilities, I like that analogy! And then the paper moves into how it models this uncertainty, which is where things get really technical. What's the next piece we need to understand about this EviLink approach?

Jane: Next, they model schema uncertainty by estimating path support across these hypotheses. They use a Beta distribution to model an item's latent selection tendency and define a credibility score based on the probability that its selection is greater than half, given the number of hypotheses and the model.

Lu: That Beta distribution modeling is sophisticated. It allows the system to quantify how much confidence it has in an item's necessity across those different paths, which is a lot of granularity for schema needs.

Meng: So, based on that score, they categorize things into sets like the required set and the uncertain set—how does that scoring translate into actual evidence acquisition?

Jane: Items with a credibility score greater than a certain threshold are placed in the required set, while those with some support but lower credibility go into the uncertain set. They then allocate evidence based on this uncertainty; items in the uncertain set get stronger evidence, like Level two semantic profiles by default, and they only pull Level three statistical profiles when the compact evidence isn't enough.

Lalam: That targeted approach to evidence acquisition seems very efficient for a large system, Meng. It stops the AI from drowning in irrelevant data while ensuring it gathers what it actually needs for those uncertain parts.

Tom: And that leads directly into the agentic refinement stage, which sounds like where the actual heavy lifting happens to resolve that uncertainty. What tools does the system use in this phase?

Jane: The refinement stage uses an agentic loop to resolve that uncertain set. The agent has read-only evidence tools, like "get field stats" or "probe sql," which it only uses when it genuinely needs stronger evidence or verification.

Lu: And the decision tools, like "keep" or "discard," are crucial there because discarding an item requires an explicit semantic reason grounded in what the item actually means. It’s not just a blind cut.

Meng: From a practical standpoint, having those read-only tools means the system is disciplined about when it seeks new information versus when it relies on what it already knows, which sounds like good resource management.

Jane: The process keeps going until the uncertain set is empty or a certain turn budget runs out, and importantly, unresolved items are kept by default to maintain recall because a correct SQL might exist along multiple plausible reasoning paths.

Lalam: It’s smart that they prioritize keeping things by default when unsure, Tom. That ensures the system doesn't accidentally cut off a valid solution just because its initial confidence score wasn't high enough.

Paper summary: Tom: So, we’ve covered the core idea—the shift to uncertainty-aware inference over multiple paths—and how they manage that with hypothesis grounding and evidence acquisition. It sounds like a significant step forward in handling the ambiguity of real-world data queries. Where do we go from here with this paper?

Jane: We're moving into the conclusions, which summarize the overall contribution of EviLink to Text-to-SQL research. The authors are focused on showing that this method offers a better balance among schema completeness, schema relevance, and token cost when compared to existing baselines like APEXSQL on complex benchmarks.

Lu: It’s interesting how they compare it against established methods; they aren't just saying their method is better in isolation, but that it manages those three competing factors simultaneously.

Meng: So, the real impact for deployment is that we get a more reliable schema context without necessarily needing the longest possible context window for every single query. That sounds like it could make Text-to-SQL systems much more scalable in production environments.

Lalam: If this works well, it means the AI can handle more nuanced business questions without breaking down under the pressure of ambiguous database structures, which really improves the utility of these tools for everyone.

Tom: Exactly. And finally, Jane, what's your take on the overall implication of this work on how we design these large language models for enterprise tasks? What's the big picture here?

Jane: The conclusion is that robust schema linking needs to move away from one-shot schema selection toward uncertainty-aware inference across multiple SQL paths, supported by acquiring evidence only where it is truly needed. EviLink achieves this by modeling path diversity and targeted evidence acquisition, successfully preserving the evidence needed for robust SQL reasoning while controlling the cost of exposing unnecessary context.

Lu: It’s about building systems that are inherently more resilient to the messiness of real data, rather than trying to force a single, perfect schema onto every query. This suggests that future Text-to-SQL models should probably be trained with this kind of path diversity in mind from the start.

Meng: I think the practical impact is that we can deploy these tools in more complex, messy enterprise databases where a single, perfect schema context simply doesn't exist. It moves Text-to-SQL closer to being truly autonomous in complex environments.

Lalam: That sounds like an AI that truly understands the nuance of a complex system, not just pattern matching on clean data. It’s about building more sophisticated reasoning capabilities into the core of how we build these helpful systems.

Tom: Wow, what a discussion! So, to wrap up our look at "EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL," we’ve seen how this paper reframes schema linking from a deterministic selection problem to an uncertainty inference one across multiple SQL paths. This seems like a solid foundation for building more resilient and contextually aware AI tools in the future.

Conclusion: Tom: So, to recap, EviLink tackles schema linking by looking across multiple possible SQL paths instead of just picking one path, using uncertainty to decide what information is truly needed and where to get it from. Jane, can you explain the main authors and what they're trying to achieve in plain language?

Jane: The paper is written by Adithya Bhaskar, Tushar Tomar, Ashutosh Sathe, and Sunita Sarawagi. Essentially, they are tackling the problem that current systems treat schema selection as a single choice when there are many valid ways to write a query.

Lu: And the core research focus is reframing this whole task from deterministic selection to uncertainty-aware inference over those different SQL paths. That's a really creative way to look at database interaction.

Meng: From a practical standpoint, that sounds like it could help us build more reliable systems when the database structure is messy and unpredictable. I wonder how this uncertainty modeling actually translates into something we can deploy without massive computational overhead.

Lalam: For me, the real vision here is that this approach could fundamentally change how we think about AI systems in enterprise settings, moving toward a culture where robustness against ambiguity is prioritized over just finding the fastest single answer.

Tom: That's a big picture, Lalam. So, when we look at the conclusion of this paper on EviLink, it boils down to what they actually proved about its title and authors?

Jane: The authors conclude that robust schema linking requires moving past single-shot selection toward uncertainty-aware inference across multiple SQL paths. They successfully show that EviLink achieves a better balance between having complete schema information, keeping the schema relevant to the query, and managing the token cost of providing all that context.

Lu: It's about proving that you can preserve enough evidence for every possible correct reasoning path while still controlling how much irrelevant data you expose. That's a really elegant constraint to manage.

Meng: I agree, Lu, managing that cost is the practical piece. If we can achieve better completeness without ballooning token usage for every single query, that makes it much more viable for real-world use cases.

Tom: So, the implication here is that EviLink gives us a framework to build AI tools that are smarter about when they need to ask for more information and when they can settle on a reasonable answer quickly. Jane, what do you see as the biggest takeaway for listeners today?

Jane: The biggest takeaway is that we should be looking at schema linking not as a simple lookup, but as an inference problem where acknowledging uncertainty and acquiring evidence selectively is the right way to go. This points toward AI systems that are much more nuanced in their understanding of complex data structures.

Lalam: I think this signals a cultural shift, Tom; we're moving toward building AI that handles real-world messiness with sophisticated reasoning, which will make these tools far more useful for everyday complex tasks.

Tom: Exactly. So, if we take all this together—the reframing of the problem and the method's results—we've seen how EviLink builds a system that is resilient and efficient in handling ambiguous database requests. This leaves us wondering what these multi-path systems can do next, Lu.

School of Software Technology, Zhejiang University

cs.CL, cs.AI

Submitted: 2026-05-28

Updated: 2026-09-28

Importance score: 81/100

The gist: Schema linking is reframed as uncertainty-aware schema-need inference over multiple plausible SQL paths, where systems distinguish required schema items from path-dependent uncertain ones and acquire

Key concepts

Multi-hypothesis Schema Grounding
This involves generating several different potential SQL queries based on the user's text. For each potential query, the system independently selects tables and fields using minimal evidence. These selections are then compared to create sets representing required, uncertain, and candidate schema elements.
Uncertainty Modeling (Beta Distribution)
The model estimates how likely an item is needed across different SQL paths using a Beta distribution. A credibility score is calculated based on this probability, determining if an item should be considered strictly required or merely uncertain.

Terminology

Summary

Schema linking is reframed as uncertainty-aware schema-need inference over multiple plausible SQL paths, where systems distinguish required schema items from path-dependent uncertain ones and acquire evidence only where needed.

Problem Reframing

The paper argues that existing methods treat schema linking as deterministic selection around a single SQL path, which limits performance in complex Text-to-SQL scenarios. The authors propose reframing the task as uncertainty-aware schema-need inference over multiple plausible SQL paths, rather than merely recover[ing] the schema appearing in one reference SQL. This reframing highlights three limiting assumptions: single-path reasoning, deterministic schema decisions, and static schema evidence. This perspective acknowledges that a question may have multiple valid SQL implementations, and since different paths require different schema elements, linking around a single path can cause irreversible information loss by discarding elements needed by other solutions.

Method Instantiation: EviLink

EviLink instantiates this reframing by combining multi-hypothesis schema grounding with uncertainty-guided evidence acquisition. The method proceeds in two stages. First, it performs multi-hypothesis schema grounding, where the model generates a set of structured hypotheses, and for each hypothesis, an independent schema grounding path selects tables and fields under lightweight evidence. These path-wise selections are then consolidated through cross-voting to initialize three sets: a required set (Sreq), an uncertain set (Sunc), and a candidate set (Scand). Second, it performs uncertainty-guided agentic refinement, focusing the agent on the uncertain set rather than the whole database.

Uncertainty Modeling and Evidence Acquisition

The paper models schema uncertainty by estimating path support across hypotheses. An item's latent selection tendency is modeled as a Beta distribution, and its credibility score is defined as "cx = Pr(px > 0.5 nx, M). This score determines set placement: items with cx ≥ τreq are placed into the required set (Sreq), while those with 0 < nx but cx < τreq" go into the uncertain set (Sunc). Evidence is allocated based on this uncertainty; schema elements in Sunc receive stronger evidence, such as Level 2 semantic profiles by default, and Level 3 statistical profiles are acquired only through get field stats when compact evidence is insufficient.

Agentic Refinement and Tool Design

The refinement stage uses an agentic loop to resolve the uncertain set. The agent utilizes specific tools to manage evidence acquisition and decisions:

  1. Evidence tools (e.g., get field stats, probe sql) are read-only, used only when needed for stronger evidence or verification.

  2. Decision tools (keep, discard) commit the refinement state, where a discard action requires an explicit semantic reason grounded in the item's semantics.

The process terminates when the uncertain set is empty or a turn budget is reached; unresolved items are kept by default to preserve recall, as stated in Rule 1: A correct SQL may exist along multiple plausible reasoning paths. Any schema item that could be needed by any path must be retained.

Systematic Validation and Results

EviLink was validated on BIRD-Dev and Spider2-Snow. On Spider2-Snow, EviLink achieved 90.15% field-level strict recall rate, used an average of 123.30K tokens, and showed improvements in downstream SQL generation under a fixed generator. The results demonstrate that EviLink achieves a better balance among schema completeness, schema relevance, and token cost compared to baselines like APEX-SQL on complex benchmarks. Furthermore, the paper shows that different accepted SQL paths can induce different gold schema annotations, supporting the necessity of preserving evidence for multiple plausible reasoning paths.

Conclusion

The work concludes that robust schema linking requires moving from one-shot schema selection to uncertainty-aware schema-need inference across multiple SQL paths, supported by evidence acquired where needed. EviLink successfully achieves this by modeling path diversity, schema uncertainty, and targeted evidence acquisition. The method is designed to preserve the evidence needed for robust SQL reasoning while controlling the cost of exposing irrelevant context.

The gist

EviLink reframes schema linking as uncertainty-aware schema-need inference over multiple plausible SQL paths, where systems distinguish required schema items from path-dependent uncertain ones and acquire evidence only where needed.

References

Anonymous. 2026. ProSPy: A Profiling-Driven SQLPython Agentic Framework for Enterprise Text-toSQL. Under review.

Adithya Bhaskar, Tushar Tomar, Ashutosh Sathe, and Sunita Sarawagi. 2023. Benchmarking and Improving Text-to-SQL Generation under Ambiguity.

Improvements for AI systems

Here are specific improvements to AI systems based on the EviLink framework, detailing what these improved systems can achieve:


  1. Improve Schema Linking by Moving from Deterministic Selection to Uncertainty-Aware Inference:

  2. Implement Multi-Hypothesis Grounding for Robust Context Preservation:

  3. Introduce Required/Uncertain Schema Bucketing for Targeted Evidence Acquisition:

  4. Utilize Hierarchical Evidence Allocation (L0 to L3) for Cost-Effective Context Building:

  5. Employ Tool-Based Agentic Refinement to Resolve Path-Dependent Uncertainty:

  6. The improved system can generate a compact schema context that is not based on a single, potentially suboptimal SQL path, but on the union of schemas required by multiple plausible SQL realizations (e.g., different join paths or filtering strategies). This prevents the irreversible information loss that occurs when schema items needed by an alternative valid query are discarded prematurely.

  7. The system can explicitly distinguish between schema elements that are consistently required across all plausible paths (the Required Set) and those whose necessity depends on which specific path is ultimately chosen (Uncertain Set). This allows the system to focus expensive evidence acquisition only where ambiguity exists, optimizing precision without incurring the token cost of uniformly exposing every detail.

  8. The system can dynamically adjust its evidence level based on schema uncertainty. It will use lightweight L0/L1 (skeleton/structure) by default but automatically request detailed L2 (semantic profiles) or L3 (statistical signals like distributions, histograms, and ranges) only for fields in the Uncertain Set. This creates a highly efficient system that avoids wasting tokens on reliably known schema items while ensuring necessary detail is available when it matters most.

  9. The system can use an agentic loop with targeted tools to verify the feasibility of its selected context. For instance, it can use a lightweight SQL probe to quickly test if a hypothesized join path is viable or if a value range overlaps, allowing it to make informed keep or discard decisions based on immediate execution feedback rather than relying solely on static profiling.

  10. The resulting Text-to-SQL system will exhibit superior end-to-end performance by achieving the best balance among schema completeness, relevance, and token cost. Specifically, it can achieve higher Strict Recall Rate (SRR) and Non-Strict Precision (NSP) on complex enterprise benchmarks (like Spider2-Snow) while significantly reducing average token usage compared to state-of-the-art baselines like APEX-SQL or LinkAlign. This leads to more accurate final SQL generation because the schema context is both sufficiently complete and clean.

Sources

Related papers