BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL

arXiv:2610.00092 · cs.CL, cs.AI, cs.DB · Submitted 2026-09-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL".

Tom: BudgetSchemaBench introduces an execution-grounded diagnostic for evaluating how database schemas fit into a model's context window when constrained by token budgets, which matters for agents working with large catalogs.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, to summarize what we’re looking at here with BudgetSchemaBench: The authors introduce this diagnostic specifically for situations where database agents have to fit schemas into a context window while facing tight token budgets. Their main claim is that this method allows us to test different schema coverage levels by sweeping through four specific budget caps across the entire catalog, keeping other factors constant.

Jane: It really matters because large catalogs can be massive, spanning many databases and thousands of columns, so agents have to make hard trade-offs between picking a few key tables or getting every little detail serialized before they run out of space. This diagnostic helps us understand that trade-off better.

Lu: What's particularly noteworthy is that they use this construction method to derive relevance labels mechanically from gold SQL without needing any human or LLM annotation for those ground truths, which is a big step in creating reliable evaluation sets.

Meng: I'm focused on the setup; they mention using a pooled eighty-database catalog and comparing three different schema representations while keeping each retriever’s table ranking fixed during the sweep, which seems like a very controlled experimental environment.

Lalam: That control is what makes it useful for us; when we can isolate variables like representation without them shifting the base retrieval performance, we get much cleaner data on where the bottlenecks are in our system's understanding of schemas.

Conclusion: Tom: Looking at the title, BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL, it really hammers home that this work provides a tool to diagnose exactly how schema context behaves when you're constrained by token limits in text-to-SQL tasks. The authors Chen Shen and Megagon Labs have developed a diagnostic that sweeps these budgets across a large catalog.

Jane: It’s important to remember the core implication is that it gives us a way to systematically probe the boundary conditions of how an AI agent handles schema selection under real cost constraints, which is something we need for scalable applications.

Lu: The work suggests that understanding this budget-swept axis isn't just academic; it points toward designing retrieval systems and model architectures that can be more robust when dealing with massive, complex data structures in production environments.

Meng: From a practical standpoint, if we can use this diagnostic to see precisely where the system fails or succeeds based on budget changes, we can make much smarter decisions about how to structure our knowledge base access layer.

Lalam: I think the real cultural impact here is showing us that rigorous, execution-grounded diagnostics are necessary for advancing AI systems; it sets a standard for evaluating these complex interactions rather than just relying on broad benchmarks.

Chen Shen

Megagon Labs

cs.CL, cs.AI, cs.DB

Submitted: 2026-09-08

Updated: 2026-09-08

Code: https://github.com/megagonlabs/budget-schema

Importance score: 89/100

The gist: BudgetSchemaBench introduces an execution-grounded diagnostic for evaluating how database schemas fit into a model's context window when constrained by token budgets, which matters for agents working

Key concepts

Schema-Context Budget K
This is the token limit allocated specifically for describing the database schema within an agent's context window. The diagnostic sweeps this budget (e.g., 2.5% to 50% of the catalog) while keeping other factors constant to measure its impact on performance.
Execution-Grounded Diagnostic
This is a testing protocol that systematically varies the schema budget K and isolates how much accuracy changes based on that variation. It uses specific probes, like removing required tables, to pinpoint exactly where the model relies on schema knowledge.
Schema-Memorization Floor
This refers to a baseline level of performance observed when required tables are intentionally removed from the context. The model still performs better than random guessing because it remembers exact table names from its training, suggesting some implicit schema knowledge remains even without explicit context.
Source-Namespace Execution Contract
This is a rule ensuring that any SQL generated must operate within the correct database environment. It checks lineage to confirm that every table referenced in the generated query belongs to the original source database.

Terminology

Summary

BudgetSchemaBench introduces an execution-grounded diagnostic for evaluating how database schemas fit into a model's context window when constrained by token budgets, which matters for agents working with large catalogs. The gist: BudgetSchemaBench is the first to combine a fixed large catalog, the budget as a controlled swept axis, and execution-grounded representation isolation (plus a source-namespace validity contract and a remove-needed-table probe).

Contribution Overview

The paper contributes three main components:

  1. An "execution-grounded diagnostic that sweeps the schema-context budget K and isolates the effect of schema representation, with derivation-first labels, a source-namespace execution contract, and a remove-the-needed-table probe."

  2. A curation method that derives relevance labels mechanically from gold SQL without human or LLM annotation and applies explicit construction-validation gates.

  3. An evaluation on a held-out complement that exercises the protocol end to end.

Diagnostic Construction and Evaluation Protocol

The diagnostic operates by sweeping four schema-context budgets (2.5%, 10%, 25%, 50% of the catalog) as absolute token caps identical across representations, while holding all other elements constant. The sweep varies schema coverage with all else held constant. Tables are packed greedily based on retriever rank order; a table is included whole if its serialization fits the remaining budget and skipped otherwise, without truncation or column pruning.

The evaluation covers three conditions:

  1. End-to-end retrieval (where the budget determines retrieval coverage).

  2. Frozen-gold (where required tables are guaranteed).

  3. A probe that removes those tables to assess sensitivity to missing schema.

Key Findings on Budget and Retrieval

Execution accuracy is found to be sensitive to the schema budget when retrieval is coverage-limited: Execution accuracy increases with the schema budget under coverage-limited retrieval. When using lexical retrieval, EX rises with K, tracking table recall (e.g., 0.60→0.90). However, under dense retrieval, the dense retriever already finds most required tables at the smallest budget, and EX remains near saturation (+0.025).

Findings on Schema Representation

The study isolates representation effects by controlling retrievers to use only table and column names (lexical name-bag-of-words ranker vs. a dense ranker based on OpenAI text-embedding-3-small), ensuring their table rankings are identical across representations. In the frozen-gold condition, representation differences are small: paired raw/enriched/hybrid differences are small (max ∆=0.02), with no contrast surviving Holm correction (every pHolm≥0.27), bounding any effect to within ±0.04 EX.

Findings on Schema Reconstruction

When the required tables are removed, the diagnostic reveals a schema-memorization floor. In 94.6% of correct predictions counted as remove-gold results, the model refers to a removed gold table by its exact name, which is interpreted as consistent with reconstruction of absent schema from parametric (weight-time) memory. This nonzero performance is termed the schema-memorization floor.

Source-Namespace Execution Contract

The protocol includes a source-namespace check that rejects queries that obtain the correct result from the wrong database. This contract requires that predicted SQL must reproduce its source-DB result, enforced by checking lineage: every base table the predicted SQL reads (across joins, sub-queries, and set operations, with CTE-local names exempted) must carry the gold database’s prefix. This check changes the classification of only 0.3% of otherwise correct end-to-end predictions.

Cross-Family Replication

The results are tested on an off-family reasoning model (DeepSeek-V4-Flash), which reproduces all three qualitative patterns, suggesting the observed patterns are not specific to the GPT family. The off-family model shows that the lexical budget axis rises by a similar margin (+0.18), while the dense axis stays nearly flat (+0.06). The remove-gold floor for this model is also low (lexical 0.01–0.05 EX, dense 0.05–0.13 EX).

Limitations

The study notes that the off-the-shelf dense retriever already achieves table recall near 1.0 at the tightest budget, suggesting this catalog may understate the difficulty of schema selection for strong retrievers. Furthermore, while representation differences are bounded in frozen-gold, they are not formally proven to be equivalent because the widest paired interval (±0.04) exceeds the preregistered δ=0.03 equivalence margin.

Improvements for AI systems

Here are the specific improvements for AI systems derived from BudgetSchemaBench, categorized by capability:


)Specific Improvements & System Capabilities

  1. Schema-Context Budgeting Ability to Handle Massive Enterprise Schemas: The system can now intelligently select and serialize relevant schema subsets from vast catalogs (thousands of columns across multiple databases) under strict token constraints (e.g., 50% budget).

  2. Retrieval-Aware Schema Selection Adaptive Budget Allocation: The system dynamically adjusts its schema selection strategy based on the retrieval mechanism used. It learns to prioritize table coverage when using lexical retrieval, while relying on dense embeddings for faster, more efficient selection when using dense retrieval, maximizing accuracy under cost constraints.

  3. Robust Schema Reconstruction (Memorization Floor) Enhanced Generalization Under Data Loss: When key tables are intentionally removed from the context (simulating budget exhaustion), the system demonstrates a quantifiable schema-memorization floor, correctly identifying and naming absent tables with high precision (94.6% exact name recall). This allows the agent to perform plausible inference or reconstruction even when explicit schema knowledge is missing, relying on parametric memory.

  4. Representation Robustness Stable Performance Across Schema Formats: The system shows remarkable resilience against variations in how the schema is represented (raw DDL vs. human-written descriptions). Under frozen-gold conditions, it maintains execution accuracy within a tight bound (±0.02 EX), indicating that the core relational structure is more critical than superficial descriptive enrichment for final SQL generation.

  5. Source-Namespace Validity Guaranteed Data Integrity: The system enforces a strict source-namespace execution contract. It will refuse to generate queries that are result-equivalent but originate from an incorrect database, preventing erroneous cross-database joins or data leakage that often plagues large catalog agents.

  6. Diagnostic Feedback Loop Automated Cost/Accuracy Monitoring: By integrating BudgetSchemaBench, system developers gain a reusable diagnostic tool to precisely measure the trade-off between schema coverage budget and execution accuracy, allowing for targeted model fine-tuning (e.g., focusing on improving dense retriever performance under tight budgets).

  7. Cross-Family Model Applicability Model Agnostic Pattern Recognition: The findings suggest that the observed sensitivity patterns (budget impact, representation effects) are not specific to a single LLM family (GPT vs. DeepSeek). This informs developers that the diagnostic is robust for optimizing various reasoning architectures when dealing with schema context.

Sources

Related papers