SDAM: Structure-Difference-Aware Memory Evolution for Complex Text-to-SQL

arXiv:2608.12338 · cs.CL · Submitted 2026-06-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SDAM: Structure-Difference-Aware Memory Evolution for Complex Text-to-SQL".

Jane: The paper was written by Keyan Xu, Dingzirui Wang, Xuanliang Zhang, Qingfu Zhu and Wanxiang Che from Harbin Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. Today we’re digging into a paper that’s got a mouthful of a title — “SDAM: Structure-Difference-Aware Memory Evolution for Complex Text-to-SQL.” Jane, I’m going to need you to unpack that for me.

Jane: Happy to, Tom. So Text-to-SQL is basically teaching a computer to turn a plain English question into a database query. Like if you ask “how many customers bought something last month,” the system writes the SQL code to find that answer. And this paper is about making that smarter over time.

Tom: Right, and the key word there is “memory.” The authors are from Harbin Institute of Technology — Keyan Xu, Dingzirui Wang, Xuanliang Zhang, Qingfu Zhu, and Wanxiang Che — and they’re saying most systems just start fresh every time. No learning from past mistakes.

Jane: Exactly. Think of it like a new employee who keeps making the same errors because nobody writes down the lessons. This paper says, let’s keep a notebook of what went wrong and what worked, and use it to do better next time.

Tom: And the “Structure-Difference-Aware” part — that’s about noticing when two different ways of writing a query actually conflict with each other. Like one path says join these tables, another says don’t. That difference is a clue something’s off.

Jane: Right, and that’s the clever bit. Instead of just picking one answer, they generate several possible SQL queries, compare the reasoning behind each one, and find where they disagree. Those disagreements become the lessons they store in memory.

Tom: So it’s not just memorizing answers, it’s memorizing the reasoning. That feels like a real step up from how most AI systems work today.

Jane: It is, and the results back it up. They improved accuracy on a hard benchmark called BIRD by two percent, and on Spider by nearly half a percent. Those numbers matter because these are tough, real-world style databases.

Tom: And that’s just the beginning. I want to get into how they actually build that memory — because that’s where the real innovation lives. Stick around.

Summary: Tom: So Jane, we’ve got the title unpacked. Now let’s talk about what this paper actually claims to achieve. The abstract lays out three big problems with existing memory systems for Text-to-SQL.

Jane: Right, and they’re pretty specific. First, weak structure analysis — most systems just record a flat list of steps, but complex SQL queries have branching logic. Second, shallow semantic understanding — they miss the deeper meaning of database fields. And third, poor schema alignment — the memory doesn’t connect well to the actual database structure.

Tom: And that third one is so practical. The paper gives a great example — a memory might say “use the name field when querying user names,” but the system might grab course.name instead of student.student name. That’s a real, everyday mistake.

Jane: Exactly. So they built SDAM to fix all three. The first fix is a reasoning tree — instead of one linear path, they generate multiple paths and merge them into a tree structure. Where the paths diverge, that’s where the interesting stuff happens.

Tom: So the tree shows the common parts and the conflicting parts. That’s how they spot structural problems.

Jane: Right. Then they add contradiction-aware reflection. They look at the tree and ask — where do these paths disagree, and why? Is it a structural conflict, a schema mismatch, or does the execution result not match what the question asked?

Tom: And then the third piece is the schema-grounded memory evolution. That’s where they actually store the lessons, but they anchor each memory to specific tables and columns in the database. So the memory knows exactly where it applies.

Jane: And they filter out bad memories too. Because if you let an AI write its own notes, some of those notes are going to be wrong. They verify each one against the schema and the existing memory bank.

Tom: So it’s a full pipeline — generate candidates, find conflicts, extract lessons, verify them, store them, and then use them on the next question. That’s a complete system, not just a trick.

Jane: It is, and the numbers show it works. They hit seventy point two percent on BIRD and eighty-eight percent on Spider with their best model. Those are strong results, especially on BIRD, which is known for being messy and hard.

Tom: I love that they don’t just claim it works — they show exactly which piece does what. That’s what we’re going to dig into next.

Improvements: Tom: Alright Jane, let’s get into the improvements this paper brings. And I want to bring in Lu and Meng for this one, because there’s a lot to chew on.

Jane: Good idea. So the paper doesn’t just say “memory is good.” It says memory needs to be structured, contradiction-driven, and schema-aware. That’s three specific improvements over the general-purpose memory systems that already exist.

Lu: And that’s the part I find genuinely exciting. General memory frameworks like ACE or ExpeL — they work for general tasks, but they’re not built for the specific challenges of SQL. SQL has structure, it has joins, it has aggregations. You can’t treat it like a generic chat task.

Tom: So you’re saying the improvement is domain-specificity. They built memory for SQL, not memory for everything.

Lu: Exactly. And the reasoning tree is the key innovation. By generating multiple schema representations — they use five different formats — they force the model to think about the same question in different ways. That diversity exposes weaknesses that a single path would miss.

Meng: From an engineering standpoint, I appreciate that they actually measured the cost. That’s something a lot of papers skip. They compared against ACE, which is a general memory framework, and they showed SDAM uses fewer LLM calls — eight point six versus ten point five per query — and less time, four point zero seconds versus four point seven.

Jane: And that’s important because memory systems can get expensive. If you’re calling the model ten times for every question, that adds up fast.

Meng: Right, and they also showed that ACE actually hurt performance on this task — it dropped accuracy by zero point seven one percent on one model. So it’s not just that SDAM is better, it’s that general-purpose memory can actively mislead a Text-to-SQL system.

Tom: That’s a strong finding. It means you can’t just bolt on any memory system and expect improvement. It has to understand the domain.

Lu: And that’s why the contradiction-aware reflection is so smart. It doesn’t just look at the final SQL — it looks at the reasoning steps, the schema choices, the execution results. It finds where the logic breaks down, and that’s where the memory gets its lessons.

Jane: And then the schema anchor — that’s what keeps the memory from drifting into hallucination. Every memory unit is tied to specific tables and columns, so when you retrieve it, you know it actually applies to the database you’re working with.

Tom: So the improvements are — structured reasoning, contradiction detection, and schema grounding. And they all work together. That’s the story of this paper.

First Page: Tom: Let’s zoom in on the first page of the paper, because there’s a figure there that really sets up the whole problem. Jane, you want to walk us through it?

Jane: Sure. The figure shows three failure cases, and each one maps to one of the three problems we’ve been talking about. The first is weak structure analysis — the system picks the wrong join or the wrong aggregation. The second is shallow semantics — a field like active user actually means the user logged in recently, not just that they’re active.

Tom: That’s the one that blew my mind. The field name is literally active user, but the real meaning is about login frequency. You can’t get that from the name alone.

Jane: Right, you have to learn it from context. And the third failure is schema alignment — the system uses the wrong column entirely, like picking course.name when it should use student.student name. That’s the kind of mistake that produces wrong answers even though the SQL runs fine.

Lu: And that’s the key insight of the whole paper — a query can execute without errors and still be completely wrong. Execution accuracy is the metric, but it’s not enough. You need semantic correctness too.

Meng: And that’s why the contradiction-aware reflection is so valuable. It catches those cases where the SQL runs but the answer doesn’t match the question. That’s the hardest error to catch because nothing crashes.

Tom: So the first page is basically saying — here are three ways systems fail, and here’s how we’re going to catch all three. That’s a really clean way to frame the research.

Jane: And it’s honest. They’re not saying “look how great our system is.” They’re saying “here are the specific problems, and here’s our solution to each one.” That’s good science.

Lu: I’d add that the figure also shows the memory evolution loop — you generate candidates, you find the contradictions, you extract the rules, and you feed them back. That loop is what makes the system improve over time without human intervention.

Tom: And that’s the part that gets me excited about where this is heading. Let’s bring in Lalam to talk about the bigger picture.

Conclusion: Tom: Alright, we’re wrapping up our discussion of “SDAM: Structure-Difference-Aware Memory Evolution for Complex Text-to-SQL.” Jane, give us the final summary.

Jane: Gladly. This paper tackles three specific problems in Text-to-SQL — weak structure analysis, shallow semantic understanding, and poor schema alignment. And it solves them with three mechanisms — a reasoning tree that exposes structural differences, contradiction-aware reflection that finds deep semantic issues, and schema-grounded memory evolution that keeps everything tied to the actual database.

Tom: And the results speak for themselves — seventy point two percent on BIRD, eighty-eight percent on Spider, and a four percent improvement over the general-purpose memory framework ACE. Plus it’s faster and cheaper.

Lu: What I find most promising is the direction this points to. The idea that memory should be domain-specific and contradiction-driven — that could apply beyond SQL. Any task where structure matters and errors are subtle could benefit from this approach.

Meng: And from a practical standpoint, the fact that they measured efficiency and showed real gains means this is something you could actually deploy. It’s not just a research curiosity.

Lalam: I’d add that this kind of self-evolving memory has cultural implications too. As AI systems get better at learning from their own mistakes, they become more reliable partners for people who aren’t database experts. That means more people can ask questions of their data and trust the answers.

Tom: That’s a beautiful way to end it. The paper is about SQL, but the bigger story is about AI that learns from its own reasoning and gets better over time.

Jane: And that’s a story worth telling. Thanks for joining us, everyone. We’ll see you on the next paper.

Tom: Take care, folks. Keep asking questions.

Keyan Xu, Dingzirui Wang, Xuanliang Zhang, Qingfu Zhu, Wanxiang Che

Harbin Institute of Technology

cs.CL

Submitted: 2026-06-03

Comments: 19 pages, 5 figures, 12tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 69/100

Key concepts

Text-to-SQL
This process teaches a computer system to translate plain English questions into functional database queries (SQL code). For example, turning 'how many customers bought something last month' into the necessary SQL statement to find that answer.
SDAM
Structure-Difference-Aware Memory Evolution. SDAM is a memory system designed for Text-to-SQL that improves accuracy by storing and utilizing lessons learned from previous queries, making it smarter over time.
Reasoning Tree
Instead of following one linear path when generating an SQL query, the system generates multiple possible paths and merges them into a tree structure. This helps expose structural differences and potential conflicts in the reasoning process.
Schema-Grounded Memory
A memory unit that is explicitly tied to specific tables and columns within a database schema. This anchoring prevents the AI from making mistakes or 'hallucinating' by ensuring stored lessons apply only where they are relevant.

Terminology

Summary

Summary

This paper introduces SDAM (Structure-Difference-Aware Memory), a structured memory evolution framework designed to improve complex Text-to-SQL generation. The authors identify three critical limitations in existing memory-based agent systems for Text-to-SQL: (i) Weak Structure Analysis, where existing methods typically only record linear reasoning processes, making it difficult to model and analyze complex SQL logic such as multi-table joins and aggregation operations; (ii) Shallow Semantic Understanding, where methods mostly remain at surface-level semantic analysis and fail to capture the deeper semantics of database fields (e.g., active user implicitly indicates login frequency exceeding a threshold); and (iii) Poor Schema Alignment, where memory lacks explicit alignment with the database schema, leading to queries inconsistent with the database structure (e.g., incorrectly using course.name instead of student.student name).

To address these issues, SDAM proposes three core mechanisms. First, the Structure-Difference Aware Reasoning Tree (§4.1.1) generates multiple schema variants and corresponding SQL candidates with reasoning paths, then topologically merge[s] semantically identical nodes across different paths while preserving structurally different or logically conflicting parts as divergent, independent branches, enabling explicit modeling of structural differences among candidate SQL queries. Second, the Contradiction-Aware Reflection (§4.1.2) leverages the reasoning tree to analyze contradictions from three perspectives: Structure Contradiction (logical conflicts among reasoning paths), Schema Contradiction (consistency between SQL and schema definitions), and Execution Contradiction (whether execution results satisfy question semantics), thereby extracting deep semantic rules from erroneous cases. Third, the Schema-Grounded Memory Evolution Mechanism (§4.1.3) performs memory extraction, filtering, and evolution, where each memory unit is defined as a quadruple m = (p, c, δ, Sanchor)—pattern, triggering condition, correction rule, and Schema Anchor that explicitly binds the memory to tables and columns in the database. The retained memories are categorized into Domain Memory, Arithmetic Memory, and SQL Syntax Memory.

The authors build SDAM-SQL, a memory evolution framework integrating SDAM into a Text-to-SQL pipeline. During memory usage (§4.2), relevant memory units are dynamically retrieved via semantic similarity (using Qwen3-Embedding-0.6B) and schema-anchor constraints (Malign = mi Sanchor(mi) ⊆ S ), then injected into SQL generation as context-aware guidance.

Experiments are conducted on Spider, BIRD, and Archer benchmarks using various open-source backbone models (Qwen3 series, Qwen2.5-Coder series, Qwen3-Coder-30B-A3B-Instruct). Key results include: (1) On BIRD-dev, SDAM-SQL with Qwen3-Coder-30B-A3B-Instruct achieves 70.2% Execution Accuracy (EX), outperforming Alpha-SQL (68.2%) by 2.0% and ExpeSQL (67.5%) by 2.7%; (2) On Spider-test, it achieves 88.0% EX, surpassing GenaSQL (87.6%) and RSL-SQL; (3) Performance gains are more pronounced on the challenging BIRD benchmark, underscoring advantages in complex SQL memory evolution and schema alignment. Across model scales, SDAM-SQL consistently improves over baselines, with gains up to 3.00% on BIRD for Qwen3-8B and peak performance of 70.21% for Qwen3-Coder-30B-A3B.

Ablation studies on BIRD show each component contributes positively: removing the reasoning tree drops performance by 3.26%, removing contradiction-aware reflection by 2.35%, removing schema anchor by 1.76%, and removing the filter by 0.46%. Difficulty-level analysis reveals that SDAM-SQL's advantage increases with task complexity, yielding gains of 2.27%, 3.67%, and 8.97% for Simple, Moderate, and Challenging settings, respectively. Comparative analysis against the general-purpose memory framework ACE shows ACE actually degrades performance (by 0.71% and 1.37% on different backbones), while SDAM-SQL improves by 3.33% and 1.17%, demonstrating domain-specific advantages. Efficiency analysis shows SDAM-SQL reduces average LLM calls per query from 10.5 to 8.6, token consumption from 30.2K to 26.6K (an 11.9% saving), and execution time from 4.7 to 4.0 seconds (a 14.9% speedup) compared to ACE.

The paper concludes that SDAM-SQL effectively alleviates structural errors, semantic deviations, and column confusion in complex SQL reasoning, and the authors plan to explore memory generalization in cross-database transfer scenarios and integrate long-cycle feedback mechanisms for further self-evolution.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:


  1. Structure-Difference Aware Reasoning Tree
  • Generate multiple schema variants (e.g., different column descriptions, table organizations) for the same database.

  • For each variant, produce a candidate SQL query and a structured reasoning path (question rephrasing → table selection → column selection → function identification → condition value identification → SQL generation).

  • Merge these paths into a unified reasoning tree, preserving shared nodes and highlighting divergent branches.

  • Use this tree to detect structural conflicts (e.g., one branch uses JOIN while another uses a subquery, or one branch applies AVG while another applies SUM).

  1. Contradiction-Aware Reflection
  • After building the reasoning tree, run a reflection step that analyzes three types of contradictions:

  • Structure Contradiction: Differences in join logic, aggregation, or filtering across branches.

  • Schema Contradiction: SQL references columns/tables not in the schema, or uses wrong data types (e.g., string vs. integer date).

  • Execution Contradiction: Execution results (e.g., COUNT=0 or NULL) conflict with the question’s intent (e.g., asking for “total” but getting a list).

  • Summarize all contradictions into a global contradiction set, which drives memory extraction.

  1. Schema-Grounded Memory Evolution
  • Extract memory units as quadruples: (pattern, condition, correction, schema anchor).

  • schema anchor explicitly binds the memory to specific tables/columns (e.g., ["table:gasstations", "column:Country", "column:Segment"]).

  • Filter out noisy or hallucinated memories by verifying against the schema and existing memory (e.g., reject a memory that says “use Status = 'Success'” if the schema shows values 1/0).

  • Categorize memories into three types:

  • Domain Memory (business rules, e.g., “active user means login frequency > threshold”)

  • Arithmetic Memory (aggregation/ranking patterns, e.g., “to find top customer, use ORDER BY SUM(consumption) DESC LIMIT 1”)

  • SQL Syntax Memory (join patterns, e.g., “when querying user names, join student on student id, not course”)

  1. Memory Retrieval and Injection
  • At inference time, compute semantic similarity between the user question and each memory unit using an embedding model (e.g., Qwen3-Embedding-0.6B).

  • Apply a schema-anchor constraint: only retrieve memories whose schema anchor is a subset of the current database schema.

  • Select the top-k most similar memories and inject them into the SQL generation prompt as structured guidance (e.g., “When querying gas stations in CZE, filter gasstations.Country = 'CZE' and gasstations.Segment = 'Premium'”).

  • Generate more accurate SQL for complex queries by explicitly modeling structural differences (e.g., multi-table joins, nested aggregations) and learning from past errors.

  • Avoid common pitfalls such as:

  • Using the wrong column (e.g., course.name instead of student.student name)

  • Misinterpreting domain semantics (e.g., treating active user as a simple status instead of a threshold condition)

  • Producing SQL that runs but returns empty results due to data type mismatches (e.g., filtering date = '2024-01-01' when the column stores 20240101)

  • Self-improve over time without ground-truth labels: each query’s reasoning tree and contradiction analysis are used to update the memory bank, so the system gets better on later, similar questions.

  • Transfer knowledge across databases when schemas share similar structures (e.g., both have a customers table with a CustomerID column), thanks to schema-anchored memory.

  • Operate efficiently with fewer LLM calls (8.6 vs. 10.5 per query) and lower token consumption (26.6K vs. 30.2K) compared to general-purpose memory frameworks, while achieving higher accuracy.

Before SDAM:

User asks: “Which customers, paying in CZK, consumed the most gas in 2011?”

System generates a broken SQL with irrelevant CASE WHEN logic and no join, returning nonsense.

After SDAM:

  • Retrieves memory: “To identify customers who consumed the most gas in 2011, aggregate consumption per customer using GROUP BY CustomerID and SUM(Consumption) from the yearmonth table, filtering for dates in 2011. Join yearmonth with customers on CustomerID and filter Currency = 'CZK'.”

  • Generates correct SQL:


SELECT T1.CustomerID

FROM yearmonth AS T1

INNER JOIN customers AS T2 ON T1.CustomerID = T2.CustomerID

WHERE T2.Currency = 'CZK' AND T1.Date BETWEEN 201101 AND 201112

GROUP BY T1.CustomerID

ORDER BY SUM(T1.Consumption) DESC

LIMIT 1

The improved AI system is a self-evolving Text-to-SQL agent that:

  • Learns from its own mistakes via structural and semantic contradiction analysis.

  • Stores reusable, schema-anchored knowledge in a structured memory bank.

  • Retrieves and applies that knowledge to new, complex queries.

  • Achieves higher execution accuracy (e.g., 70.2% on BIRD, 88.0% on Spider) with lower computational cost than existing memory-based methods.

Sources

Related papers