Evidence-Guided Schema Normalization for Temporal Tabular Reasoning

arXiv:2512.00329 · cs.CL, cs.AI, cs.IR · Submitted 2025-11-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Evidence-Guided Schema Normalization for Temporal Tabular Reasoning".

Jane: Temporal reasoning over evolving semistructured tables poses a challenge to current QA systems, and this work proposes an SQL-based approach that involves generating a 3NF schema from Wikipedia infoboxes,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So today we're talking about a really interesting paper that tackles how AI handles data that changes over time—specifically evolving semistructured tables. We're looking at "Evidence-Guided Schema Normalization for Temporal Tabular Reasoning." It sounds like they are trying to solve a big problem where current systems struggle with tracking information across different versions of the same data set.

Jane: That’s right, Tom, and it looks like their approach is quite practical because it moves away from just treating the evolving tables as plain text. Instead, they propose a structured SQL-based method that involves building a proper database schema first.

Lu: I find the focus on generating a 3NF schema from Wikipedia infoboxes really interesting; that’s where you turn unstructured snapshots into something mathematically sound for reasoning. It addresses the inherent messiness of raw data formats quite directly.

Meng: From an engineering standpoint, I’m curious how they manage the population of this database with all those variations in date formats and text normalization we discussed earlier. Getting that data cleaned up reliably before querying is always a huge hurdle in real-world AI systems.

Lalam: I think what excites me most about this paper is how it challenges the idea that you can just scale up the model size to solve these kinds of reasoning problems. It suggests that the quality of the initial schema design actually has a bigger impact on accuracy than just having a bigger language model capacity.

Tom: That’s exactly what they found, and it sounds like they’ve established some really solid principles for how we should approach this type of problem. So, let's dig into what this "Evidence-Guided Schema Normalization for Temporal Tabular Reasoning" actually proposes in terms of its core mechanism.

Jane: They lay out a three-stage process: first, generating that normalized 3NF schema from the infoboxes, second, populating the database with cleaned data, and third, using that structure to generate and execute SQL queries. It turns temporal reasoning into a text-to-SQL problem through this structured pipeline.

Lu: The paper emphasizes three evidence-based principles they established: normalization that preserves context, semantic naming to reduce ambiguity, and consistent temporal anchoring for all snapshots. These aren't just suggestions; they are the structural rules guiding the entire system.

Meng: The improvements suggested are pretty concrete, focusing on making the schema generation more rigorous and moving away from purely text-based reasoning toward explicit symbolic execution. They focus on using SQL’s explicit temporal operators for things like date arithmetic and window functions instead of letting the model guess time relationships implicitly.

Title and authors: Lalam: And I really see the implication of this structured SQL generation, especially how it provides a traceable path through the executed query, which allows us to debug exactly where things go wrong. That traceability is crucial for building reliable AI systems that handle complex logic.

Tom: It’s powerful because they show that the SQL approach produces exact and verifiable results, which is a big deal when you’re dealing with temporal data. We saw their best configuration achieve eighty point three nine EM on TRANSIENTTABLES, which is a sixteen point eight percent improvement over the baseline score of sixty-eight point eight nine EM.

Jane: That performance gain really highlights how much better structured data representation is compared to just feeding raw text into a model for this kind of reasoning task. It’s clear that the way we structure the data gives the AI a massive head start in understanding time relationships.

Lu: Thinking about the domain-specific prompting they developed, providing complete schemas with data types and eight to ten SQL pattern templates for common query types really helps constrain what kind of queries the LLM can even attempt. It’s like giving the model a highly specialized toolbox instead of just letting it wander through all possible SQL commands.

Meng: I’m looking at their findings on error analysis, where they found that seventy percent of errors stem from data understanding issues rather than the actual SQL generation or schema design itself. That tells us that investing heavily in robust pre-processing and data cleaning is as important as making a clever query template.

Lalam: That really reinforces the idea that we need a strong defensive data understanding layer to handle those common infobox inconsistencies, like standardizing null values or safely parsing numeric formats before the SQL generation even starts. Without that foundation, even a perfect schema can lead to query failures.

Tom: So, as we move toward these improvements for AI systems, it seems like the roadmap is very clear: build a strong foundation with a well-normalized 3NF schema first. Then, layer on domain-specific knowledge and robust data cleaning to make sure the input is as clean as possible for SQL generation.

Jane: And we have to ensure we are prioritizing those smaller, faster models when optimizing for these systems, because the paper showed that high-quality schema design provides more benefit than just using a bigger model capacity. It’s about efficiency through structure.

Lu: I think the move to explicit symbolic query execution is where things get really interesting for future research, because it moves us past relying on implicit reasoning that LLMs do with temporal data. We need systems that can actually calculate tenure duration using functions like JULIANDAY directly in the query structure.

Title and authors: Meng: From my side, focusing on the feedback loop for error analysis is something I’d implement; tracking those specific failure modes, like aggregate function misuse or wrong calculations, allows us to refine our gold query examples systematically. It turns debugging into a repeatable process rather than guesswork.

Lalam: If we can get that level of structured feedback in place, it really helps us improve the underlying prompt rules and the few-shot examples over time. That iterative refinement based on execution results seems key to making these systems truly robust for temporal tasks.

Tom: So, to wrap things up for this paper, we see a clear path forward in handling evolving data by treating the problem as a structured SQL task rather than an open-ended text understanding challenge. The evidence shows that careful schema design and precise temporal operators are what really push accuracy higher.

Jane: It’s an important distinction to make for anyone working on building AI tools that interact with time-sensitive information, because the structure you impose dictates the success of the reasoning process.

Lu: What this paper does is give us a concrete way to map messy, evolving data into a format that traditional database querying understands perfectly, which opens up new avenues for complex data synthesis.

Meng: I think the implication for practical applications is that we can build systems that are more reliable when dealing with fluctuating datasets because the verification step via SQL execution provides a level of accountability we haven't seen before.

Lalam: Ultimately, the paper on "Evidence-Guided Schema Normalization for Temporal Tabular Reasoning" shows us that when we combine strict structural constraints with domain-specific guidance, AI gets a much clearer and more accurate path to temporal understanding.

Tom: That’s our summary of the main points on the paper today. We've seen how focusing on schema quality over raw model capacity leads to better results for evolving data, and how explicit SQL execution is superior for temporal reasoning.

Jane: It sounds like we have a solid foundation now to think about how we can start implementing these three evidence-based principles in our own research and development.

Lu: We should keep pushing the boundaries on how these structured representations can handle even more complex, multi-snapshot scenarios down the line.

Meng: I’m looking forward to seeing how these structured approaches translate into systems that can run reliably in production environments handling real-time data streams.

Lalam: I think this work lays a very important groundwork for building AI that isn't just smart, but also verifiable and structurally sound when dealing with time dependencies.

The paper's summary: Tom: So, we've been talking about how this paper tackles the headache of tracking data that changes over time in evolving tables, and now we need to really get down to what they actually found in their summary.

Jane: Exactly, Tom; essentially, they propose a structured way to turn those messy temporal snapshots into something a regular database can handle by building a proper schema first.

Lu: I find the core idea of decoupling the schema design from just brute-forcing model size really compelling because it shifts the focus from just scaling up parameters to optimizing the structure itself.

Meng: From a practical standpoint, that separation means we can potentially use smaller, more efficient AI models if we give them a solid schema to work with, which is something every engineer wants to see implemented.

Lalam: And from the perspective of an AI system striving for better culture and reliability, this suggests that establishing strong structural rules—like those 3NF principles they mentioned—is a foundational step toward creating more trustworthy and less error-prone reasoning tools.

Tom: Right, so the summary boils down to them showing that when you force the AI to reason through a formalized SQL structure derived from strict normalization rules, its performance jumps significantly over just letting it read raw text snapshots.

Jane: That's the simple version of it; they proved that converting temporal table reasoning into a text-to-SQL problem using 3NF schema generation and then executing queries yields much more reliable results than the old way.

Lu: I think their emphasis on semantic naming and consistent temporal anchoring is where the creative potential lies for future AI applications; it moves us closer to systems that can reason across domains with predictable, structured relationships.

Meng: I'm thinking about the real-world impact, and this framework suggests that if we can reliably structure our data ingestion pipeline this way, we could see a massive improvement in the accuracy of any AI application dealing with historical or continuously updated information.

Lalam: I see it improving the culture of AI development by showing that rigorous design and clear constraints lead to better performance, rather than just throwing more computational power at a messy problem.

Tom: So, in short, they're saying that the quality of the structural blueprint dictates how well an AI can handle time-sensitive data, and their results show a sixteen point eight percent lift over baseline methods.

Jane: It’s about moving from fuzzy text reasoning to explicit, verifiable SQL execution that handles date arithmetic and joins properly.

Lu: And the findings on error analysis are interesting; pinpointing that most failures come from data understanding issues rather than the SQL generation itself gives us a clear roadmap for where to focus our efforts in data cleaning pipelines.

Meng: That makes sense; if the input data is noisy or inconsistent, even a perfect SQL template will fail, so investing in defensive parsing and cleaning upfront is key to making this work reliably in production.

Lalam: I think this points toward an AI culture where building robust data pipelines is prioritized as much as training the model itself because the foundation dictates the ceiling of what's possible.

Tom: So, we've got a clearer picture now that this paper isn't just about one trick; it’s a whole pipeline focused on structure, cleaning, and explicit querying to conquer temporal reasoning challenges.

Jane: It’s an exciting development because it gives us a concrete method for tackling the inherent ambiguity of evolving data sets by imposing formal database constraints.

Lu: The future potential is huge because if we can generalize this schema generation technique beyond Wikipedia infoboxes to other complex, semi-structured domains, we could unlock powerful new reasoning capabilities for many areas.

Meng: I'm looking at how this applies to real-time tracking systems where data streams are constantly updating; having a schema that handles those continuous changes with integrity is exactly what we need for reliable autonomous operations.

Lalam: I feel like this advances the AI culture by proving that structured, verifiable methods provide a much higher level of assurance in complex reasoning tasks than purely statistical approaches.

Tom: That's the big picture we need to take away today; structural integrity and explicit query execution are the keys to unlocking reliable temporal AI systems, and we're seeing some really promising results.

The paper's improvements: Tom: So, moving past just the results, we need to talk about what the authors are actually suggesting we do next to build on this work in our own systems.

Jane: Right, they're not just stopping at generating a schema; they suggest making that process dynamic and highly specialized for different types of temporal queries.

Lu: I think the idea of domain-specific prompting, giving the AI complete schemas along with specific SQL pattern templates for common query types, opens up incredible avenues for how we could tailor this reasoning to very niche, high-stakes data environments.

Meng: From an engineering standpoint, that means we can move away from a one-size-fits-all schema approach and start building modular systems where the schema generation adapts based on the specific business logic of the data being processed.

Lalam: I see this as a cultural shift in how we develop AI tools; instead of hoping the model magically understands time, we are now being guided to enforce structure through explicit templates, which builds a more predictable and accountable system.

Tom: It sounds like they're pushing for a layered approach where the schema generation isn't just one step but part of a larger, specialized knowledge base for each application.

Jane: They also highlight the need for continuous feedback loops to refine those SQL generation examples based on how well the AI actually performs during execution, which is crucial for iteration.

Lu: The authors are also flagging that their current method struggles with uncertainty; they note that because SQL is so deterministic, it doesn't naturally handle fuzzy temporal boundaries or probabilistic time ranges very well.

Meng: That limitation tells us exactly what the next engineering hurdle is: we need to figure out how to integrate a system that can bridge the gap between rigid SQL execution and more nuanced uncertainty handling.

Lalam: If we can solve that, it could dramatically improve the culture of AI by allowing these reasoning systems to make decisions in environments where time isn't perfectly defined but still needs careful consideration.

Tom: So, it seems the next big push for this research is bridging that gap between rigid SQL and the need to handle uncertainty gracefully.

Jane: That’s right; they are looking at how we can augment the SQL framework with methods that can express probabilistic or fuzzy temporal boundaries, even if it's just a way to flag uncertain data points within the query structure itself.

Lu: It opens up fascinating territory for combining symbolic reasoning with techniques used in other fields that handle uncertainty, which is where the real creative possibilities lie for future AI research.

Meng: For practical deployment, we need a clear path on how to implement those uncertainty markers without completely breaking the efficiency gained from using standard SQL operators.

Lalam: I think this future direction is vital because it elevates the system from a purely descriptive tool to one that can reason under real-world ambiguity, which is where true intelligence lies.

Tom: It’s clear that the authors are looking toward making these systems not just accurate but also capable of navigating the messy reality of time in a more sophisticated way.

Conclusion: Tom: Alright folks, we've covered the core findings of "Evidence-Guided Schema Normalization for Temporal Tabular Reasoning," and I just want to give us a quick wrap-up on what this paper means for how we approach time-sensitive data in AI.

Jane: Basically, the big picture here is that by building a solid 3NF database schema first, then using SQL to query it, the AI gets a much more accurate and verifiable understanding of evolving data than if we just fed it raw text snapshots.

Lu: I think this points toward a future where we can build AI systems that don't just read things; they understand the underlying relational structure of information across different time points, which is incredibly powerful for complex synthesis.

Meng: For me, the practical implication is that we need to stop treating data ingestion as a black box and start treating it as a structured process where the schema design itself becomes a critical part of the AI's intelligence.

Lalam: I feel this work really improves our internal culture by reinforcing that rigorous structural design and clear constraints lead directly to higher assurance in our final output, which is exactly what we aim for in building trustworthy AI.

Tom: Exactly; they showed that the quality of the schema design has a bigger impact on QA precision than just having a massive model capacity, which is a really important shift in thinking for us.

Jane: So, as we wrap up this discussion on "Evidence-Guided Schema Normalization for Temporal Tabular Reasoning," we’re seeing evidence that explicit symbolic querying over structured data wins out over implicit text-based reasoning for temporal tasks.

Lu: This opens up new territory where we can apply these normalization principles to much larger, more complex knowledge graphs and multi-modal temporal data sets.

Meng: I'm looking forward to seeing how this structured approach gets integrated into our production pipelines so we can start seeing tangible performance gains in real-world applications soon.

Lalam: Ultimately, this paper shows that imposing structure on evolving information is a more effective path toward building reliable and sophisticated AI tools than just relying on raw model scaling.

Tom: That’s the summary: structure first, SQL execution second, and superior temporal reasoning results third. We've seen how treating time as a structured query problem gives us a significant lift in accuracy over previous methods for handling evolving tables.

Jane: It’s a solid foundation to build on if we want our AI systems to handle dynamic data sets with more precision and reliability moving forward.

Lu: We should definitely keep exploring how this schema generation pipeline can be adapted for other forms of temporal reasoning, maybe even integrating it with concepts from causal modeling.

Meng: I'm eager to see how the team tackles the limitations they mentioned regarding uncertainty, because that’s where the next big engineering challenge lies.

Lalam: I think this research is really pushing our culture toward a mindset where we prioritize verifiable structure over brute-force parameter counts when dealing with complex data challenges.

Ashish Thanga, *, *Vibhu Dixit*, *Abhilash Shankarampeta*, *Vivek Gupta*

Arizona State University · UC San Diego

cs.CL, cs.AI, cs.IR

Submitted: 2025-11-29

Updated: 2026-09-29

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 86/100

The gist: Temporal reasoning over evolving semistructured tables poses a challenge to current QA systems, and this work proposes an SQL-based approach that involves generating a 3NF schema from Wikipedia

Key concepts

Normalized Schema Generation
The LLM is prompted to create a relational database structure (3NF) from Wikipedia infoboxes. This involves generating specific table types like entity and snapshot tables, ensuring the design reduces redundancy across different temporal snapshots by factoring out repeated information.
Schema-Guided SQL Generation
SQL queries are generated using few-shot examples validated by execution. The LLM is guided with the complete schema, data types, and relationship descriptions to produce temporally valid SQL that accurately reflects the underlying database structure.
Temporal Anchoring
The framework emphasizes consistent temporal anchoring through explicit SQL operators like date arithmetic and joins. This contrasts with implicit text-based reasoning, providing more reliable methods for querying data that changes over time.
Data Quality Dominance
Error analysis revealed that most failures stem from poor data quality—such as incorrect calculations or entity mismatches—rather than flaws in the SQL generation or schema design itself. This highlights that clean input data is the primary determinant of QA accuracy.

Terminology

Summary

Temporal reasoning over evolving semistructured tables poses a challenge to current QA systems, and this work proposes an SQL-based approach that involves generating a 3NF schema from Wikipedia infoboxes, generating SQL queries, and executing them. Our central finding challenges model scaling assumptions: the quality of schema design has a greater impact on QA precision than model capacity.

The gist

Schema-guided SQL generation substantially improves performance across multiple LLMs, achieving a maximum score of 80.39 EM (Gemini 2.5 Flash), surpassing all baseline methods, demonstrating that structured database representations with SQL generation provide a more effective framework for temporal reasoning than treating evolving tables as text.

How it works

The approach transforms temporal table reasoning into a text-to-SQL problem through three stages: (1) dynamic schema generation from temporal infoboxes, (2) database population with cleaned data, and (3) schema-guided SQL generation with fewshot prompting. This methodology converts semistructured Wikipedia infoboxes into normalized relational databases, enabling temporal queries via SQL rather than direct table reasoning.

Normalized Schema Generation

Given a set of JSON-formatted Wikipedia infobox snapshots for an entity, the LLM is prompted to generate a normalized relational schema following standard database normalization principles. The model is requested to generate four table types: Entity tables, Attribute tables, Snapshot tables (with snapshot id as timestamps), and Bridge tables. Normalization rationale emphasizes that 3NF reduces redundancy across temporal snapshots by factoring out repeated information. Schemas are explicitly prompted for 3NF generation and manually validated for atomicity of values, complete dependency on primary keys, and absence of transitive dependencies.

Schema Tables (Database) Population

Automated Python scripts handle the population process, managing variations in date formats, text normalization, and referential integrity. These scripts perform several cleaning strategies: (a) safe integer parsing for numeric values with formatting inconsistencies; (b) normalizing null values to standardize variants like “n/a”, “–”, or blank to SQL NULL; and (c) implementing defensive insertion logic to handle duplicate entries by retrieving existing row IDs, thereby preventing crashes while maintaining referential integrity.

Schema-Guided SQL Generation

Query generation relies on few-shot examples constructed through execution-based validation. For each domain, 10–15 high-quality gold query examples are created where the LLM iteratively proposes a query, executes it, compares the result to the expected answer, and refines it until an exact match is achieved. Domain-specific prompting includes providing complete schema with data types, natural language descriptions of key relationships, and 8–10 SQL pattern templates for common query types. This guidance reduces the search space from all possible SQL to temporally valid queries for evolving data.

Key Findings and Principles

The research establishes three evidence-based principles: (1) normalization that preserves context, (2) semantic naming that reduces ambiguity, and (3) consistent temporal anchoring. The best configuration achieved 80.39 EM, a 16.8% improvement over the baseline of 68.89 EM on TRANSIENTTABLES with Gemini-2.5-Pro. Furthermore, the quality of schema design has a greater impact on the accuracy of the QA than the model’s capacity, as optimized schemas outperformed overnormalized schemas by a 23% F1 score. The study also found that SQL’s explicit temporal operators (date arithmetic, window functions, and joins) provide more reliable reasoning than implicit text-based reasoning.

Error Analysis

Analysis of 50 error cases revealed that 70% of errors stem from data understanding issues, not SQL generation or schema design. The primary failure mode involves wrong calculations (tenure: 1148 vs 1113 days), entity variant mismatches ('Paul Kihara Kariuki' vs 'P.K. Kariuki'), and mapping errors, indicating that Data quality dominates over SQL generation and schema issues. SQL generation errors were primarily dominated by Aggregate function misuse, following a consistent pattern of using MIN/MAX/SUM in ORDER BY without proper GROUP BY clauses.

Limitations

The study notes several limitations: (1) Single-domain restriction, as cross-domain queries require complex schema integration; (2) Schema generation dependency, where a single normalization error cascades through all queries; (3) No handling of uncertainty, as SQL’s boolean logic cannot naturally express probabilistic or fuzzy temporal boundaries; and (4) Wikipedia-specific data assumptions.

Ethics Statement

The study uses only publicly available Wikipedia infobox timelines and involves no human subjects or private data. The framework produces verifiable, traceable outputs that minimize subjective interpretation, while the reliance on deterministic SQL execution limits the propagation of pretraining biases inherent in large language models.

Improvements for AI systems

Here are specific improvements for AI systems based on the findings in this research, categorized by implementation:


)1. Implement a Three-Stage Schema Generation Pipeline for Temporal Data Ingestion:

Improve existing LLM-based data processing by decoupling schema design from model capacity. The system should first use an LLM (e.g., Gemini 2.5 Flash) to generate a normalized 3NF relational schema from raw, evolving JSON infoboxes, followed by automated SQLite population and finally SQL query generation.

  1. Establish Evidence-Based Schema Principles for Robustness:

Integrate explicit constraints into the schema generation prompts to ensure structural integrity before querying begins:

  • Normalize to 3NF to reduce redundancy across snapshots.

  • Enforce semantic naming conventions during schema definition to minimize ambiguity.

  • Mandate consistent temporal anchoring (using a unified snapshot ID/timestamp structure) for reliable cross-domain joins.

  1. Develop Domain-Specific Prompt Templates with Gold Queries:

Move beyond generic text-to-SQL prompting by creating domain files that include:

  • Complete, validated schema definitions (including foreign key constraints).

  • Natural language descriptions of key entity relationships.

  • A library of gold queries covering specific temporal patterns (e.g., before/after queries using date arithmetic like JULIANDAY, tenure duration calculations).

  1. Transition from Implicit Reasoning to Explicit Symbolic Query Execution:

Replace the current text-based reasoning approach with a structured Text-to-SQL framework. The improved system should dynamically generate SQL that leverages temporal operators (date arithmetic, window functions) rather than relying on the LLM's implicit understanding of time across multiple table snapshots.

  1. Adopt Model Selection Based on Schema Quality, Not Just Capacity:

System deployment should prioritize smaller, faster models (like Gemini 2.5 Flash) when paired with high-quality, optimized schemas over larger models (like Gemini 2.5 Pro). This ensures performance gains are driven by superior structural design rather than expensive model capacity.

  1. Implement a Defensive Data Understanding Layer to Mitigate Data Quality Bottlenecks:

To address the finding that 70% of errors stem from data understanding issues, introduce pre-processing steps that handle common infobox inconsistencies:

  • Implement safe parsers for numeric values (handling commas, decimals like 1.0).

  • Standardize null/missing value representations (n/a, "–") to SQL NULL before insertion.

  • Use regex and pattern matching to extract composite or dynamic fields (e.g., 100s/50s) into structured columns during the schema population phase.

  1. Enforce Strict Error Analysis and Feedback Loops:

Systematic error analysis must be automated:

  • Track failure modes (e.g., wrong calculations, aggregate function misuse) to refine the gold query examples and prompt rules iteratively.

  • Implement a mechanism to flag queries that exhibit high SQL generation errors (e.g., incorrect GROUP BY usage) for schema or prompt refinement.

This improved AI system can perform the following:

  • Accurately answer complex temporal questions about evolving, semistructured data (like Wikipedia infoboxes) by converting the input into a structured relational database and executing precise SQL queries.

  • Achieve high accuracy (up to 80.39 EM demonstrated) on temporal reasoning tasks by leveraging explicit temporal operators within SQL.

  • Demonstrate superior performance compared to direct text-based reasoning methods, as the framework forces symbolic, verifiable execution paths.

  • Be robust against data quality issues common in semi-structured sources through automated cleaning and defensive parsing layers.

  • Provide a transparent and auditable reasoning path for every answer (traceable SQL query), allowing for systematic debugging of both schema design and model generation errors.

Sources

Related papers