Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models

arXiv:2607.21636 · cs.LG, cs.AI · Submitted 2026-07-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models".

Jane: The paper was written by Jie Zhang from Accenture, Tokyo 107-8672, Japan and Accenture.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 1 — Tom and Jane discuss title and authors of the paper 'Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models' and its implications. Explain in simple terms; do not repeat what earlier segments covered.: Tom: We’ve talked about how crucial inter-column dependency is, but the paper "Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models" really focuses on why our current methods fail to see that structure.

Jane: It seems like many standard benchmarks are simply blind to this relational structure, which is a huge problem when we're trying to ensure data quality in fields like fraud detection or clinical risk assessment.

Lu: The title itself is so revealing because it suggests that the gap isn't just some random measurement error; it’s a measurable deficit in the fidelity of the dependency between columns.

Meng: From an engineering standpoint, this implies that we need to rethink how we define "successful" data generation, moving beyond surface-level visual similarity to something much more rigorous.

Lalam: We have a responsibility to understand these flaws before deploying systems that could impact people’s lives based on flawed synthetic inputs. This paper is a mirror showing us where our current testing methods are insufficient.

Tom: The authors are setting the stage by showing how little the existing linear tests—the C2ST and Trend scores—can see, which is really alarming when you think of how much data we rely on.

Jane: It’s a good reminder that just having marginal distributions match isn' not enough; they need to co-occur in the right ways too, otherwise the model is fundamentally incomplete.

Lu: This work suggests that the next generation of models needs to understand dependency as a core architectural requirement, rather than an optional feature.

Meng: It forces us to ask questions about how our current pipelines are validated—are we just checking if they look like real data, or are we checking if they *act* like real data?

Lalam: The implications for building trustworthy AI systems are massive; we can' demanding proof of functional correctness.

Tom: It’s a conversation that needs to be happening more I think, about the genuine structural integrity of how synthetic data behaves.

Paper discussion segment 2 — Tom and Jane discuss the paper's summary of the paper 'Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models' and its implications. Explain in simple terms; do not repeat what earlier segments covered.: Tom: So, after establishing that existing metrics are blind, the authors introduced a powerful diagnostic tool by splitting the C2ST score into three components: marginal, dependency (`dep`), and cross terms (`depcross`).

Jane: This decomposition is incredibly useful because it allows us to isolate exactly what's wrong with the generator—we can see if it's missing structure or if its marginal distributions are off.

Lu: The way they define `dep` as a contrast between this score and a fully factorized reference provides such clarity; we’ are no longer just observing a vague failure, we are quantifying the structural shortfall.

Meng: I like that this framework allows us to pinpoint the specific failures in practical applications; instead of saying, "The data looks bad," we can say, "The conditional probability between Feature A and Feature B is broken."

Lalam: This moves us toward a culture of verifiable accountability in AI development, allowing regulatory bodies to understand exactly where a generative model needs to be improved.

Tom: The results on TabbyFlow and TabDiff show this gap is persistent across different datasets, which is a very strong indication that it's not just an issue with one specific data set.

Jane: It’s consistent across the benchmarks, showing that the problem isn't isolated; it’s a systemic weakness in how these models are trained to capture relationships.

Lu: This consistency tells us that regardless of the model architecture, there is a fundamental challenge in capturing joint structures without explicit supervision.

Meng: It means we can apply this diagnostic to any production-level generator and get reliable feedback on its structural integrity, which is a massive win for automation.

Lalam: This is the right level of detail needed to ensure that trust in AI is based on genuine proof rather than just superficial resemblance.

Tom: We’re getting much closer to understanding exactly what these models are failing to learn compared to real data, and it seems like we have a roadmap for the next steps.

Paper discussion segment 3 — Tom and Jane discuss the improvements the paper suggests of the paper 'Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models' and its implications. Explain in simple terms; do not repeat what earlier segments covered.: Tom: We’ve seen that this gap exists, but it’s not just a random shortfall; it's a measurable component of the model's performance compared to the real data oracle.

Jane: The paper strongly suggests that since dependency is vital for minority-class utility—like finding rare fraud—we can’t ignore this structural deficit even though current metrics do.

Lu: I think the theoretical insight here is that the models are capable of representing joint structure, but they aren't being pushed to do so by the objective function itself.

Meng: The practical implication for us is that we need a new training signal; instead of just hoping it works, we should be actively rewarding models for capturing those inter-column relationships.

Lalam: This suggests a shift in our development culture—we’ must build systems where the AI actively seeks out and learns these hidden dependencies to create trustworthy outputs.

Tom: The authors are very clear that because this gap is so substantial, we cannot dismiss it as just some minor artifact of sampling or noise.

Jane: It's a tangible, measurable failure point that has a real impact on downstream applications like predictive analytics and risk management.

Lu: This work shows that the next generation of models need to understand the interplay between columns as a foundational element of their learning process.

Meng: We are looking at specific interventions—modifying the objective or adding extra capacity—to see if we can close this gap, which is a very focused engineering path forward.

Lalam: By highlighting that dependency matters, they' are giving us the justification to demand more rigorous training protocols for our AI systems.

Tom: It’s moving the conversation toward designing models that *must* learn the joint behavior of a stronger case for verifiable operational logic.

Conclusion — Tom and Jane lead the wrap-up: they summarize the paper' implications and say goodbye to it, getting ready for the next paper. Before the goodbye, Lu, Meng, Lalam each gets one final short turn to weigh in.: Tom: So, we’ve walked through "Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models" and seen how it's not enough for synthetic data to just look pretty.

Jane: We've learned that structural integrity is a mandatory requirement, not an optional extra, because without it, the utility of our AI systems is severely compromised.

Lu: The next generation of models must be built with dependency awareness baked into their core architecture, making it a prerequisite for any meaningful results.

Meng: For us in the industry, this provides a clear path: we have actionable metrics to enforce structural compliance in real-world deployments.

Lalam: This allows us to build genuine audit trails into our AI systems—proving the data's behavior under pressure, not just claiming it looks right.

Tom: That brings us to the practical implications of demanding a verifiable operational logic from any generative model.

Jane: We are genuinely excited to apply these rigorous new standards as we move on to the next paper, knowing exactly what we need to look out for in every single one.

Lu: I hope that this diagnostic tool becomes widely adopted so that the entire field moving forward will be architecturally sound, rather than relying on patch-up solutions for existing models.

Meng: It really does set a new, higher bar—a bar rooted in verifiable operational logic that synthetic data must adhere to in the real world.

Lalam: This whole discussion is a crucial moment for the field; "Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models" gives us both the language and the tools to elevate our entire conversation around data trust.

Tom: Thank you all for this deep dive into structural fidelity. We've truly gained a much deeper appreciation for what it means to achieve true, verifiable accuracy in AI outputs.

Jie Zhang

Accenture, Tokyo 107-8672, Japan · Accenture

cs.LG, cs.AI

Submitted: 2026-07-20

Updated: 2026-08-25

Importance score: 80/100

The gist: * Motivation and Problem Definition The utility of synthetic tabular data relies not only on preserving marginal distributions but also on capturing "inter-column dependency, a requirement

Key concepts

Inter-Column Dependency Gap
This refers to the measurable deficit or lack of fidelity in the relationship between different columns within a dataset. Current testing methods are often blind to this structural shortfall, making it difficult to ensure data quality in real-world applications.
C2ST Decomposition
The authors propose a diagnostic tool that splits the existing C2ST score into three parts: marginal, dependency (dep), and cross terms (depcross). This framework allows researchers to isolate exactly where a generative model is failing structurally.
Functional Correctness vs. Visual Similarity
The discussion emphasizes that synthetic data must be tested not just for how it looks like real data, but for how it acts like real data. This requires moving beyond surface-level visual resemblance to achieve genuine structural integrity.

Terminology

Summary

** Motivation and Problem Definition**

The utility of synthetic tabular data relies not only on preserving marginal distributions but also on capturing inter-column dependency, a requirement particularly critical in financial fraud detection... and clinical risk prognoses. Preserving this inter-column dependency is deemed pivotal to the overall fidelity of synthetic tabular data and its downstream utility, particularly for minority-class detection.

However, standard evaluation metrics are largely blind to this structure. The paper notes that a "fullyfactorized baseline that destroys all inter-column dependency still appears nearly real under the commonly reported linear classifier two-sample test (C2ST)—a known weakness we confirm on four benchmarks—and is only mildly penalized by pairwise Trend scores. This suggests that existing metrics primarily certify marginals rather than dependency, rendering it largely blind to inter-column structure."

** The Dependency-Aware Diagnostic Tool**

To address this limitation, the authors formulated a diagnostic tool that decomposes the C2ST score. By equipping C2ST with an interaction-sensitive learner (XGB-C2ST), they decompose the score into three components: marginal, dependency, and numerical–categorical cross terms.

The methodology involves comparing these components against two fixed references:

  1. The Fully-Factorized Reference (FF): This reference preserves marginal[s] but possesses zero dependency and serves as a low-fidelity baseline where the dependency component spans the entire deficit.

  2. The Oracle: This is a fresh subsample of real training data, serving as a high-fidelity reference (where C2ST is near chance, or about 0.5).

The decomposition allows researchers to localize the deficit inside one discriminator’s score and measure mixed numerical–categorical coupling, which traditional cumulant-based comparisons cannot achieve.

** Empirical Findings: The Persistent Dependency Gap**

Applying this diagnostic to representative generators—TabbyFlow/EF-VFM (flow matching) and TabDiff (diffusion)—the authors found a persistent dependency gap of comparable magnitude in both.

Key findings regarding the gap include:

  • Necessity for Utility: Dependency is necessary for minority-class utility—a zero-dependency reference collapses it, indicating that destroying inter-column structure severely harms downstream tasks.

  • The Residual Shortfall: Despite the critical nature of dependency, the generators’ residual shortfalls, however, are small and do not track the measured gap across datasets. This means the operational difference in minority-class F1 between a generator and the real data oracle is often minimal compared to how much structure was lost.

** Analysis of Potential Explanations**

The paper investigated several potential causes for this gap:

  1. Structural Limitation (H1): The hypothesis that a mean-field (column-factorized) objective cannot represent inter-column dependency was ruled out, as the mean-field optimum recovers the true per-dimension posterior means... in the infinite-capacity limit.

  2. Insufficient Capacity (H2): The hypothesis that the baseline lacks the capacity to capture dependency was tested via a width sweep (scaling up to 16 times capacity). The results showed no significant change in the dependency component on datasets where training was clean, leading the authors to conclude that a 16× increase does not close the gap.

  3. Optimization/Objective Limitation (H3): The authors argue that the deficit plausibly lies in what the objective rewards, rather than in what the architecture can represent or how large it is. This suggests a lack of gradient pressure toward joint fidelity at finite capacity.

** Failure of Remedial Interventions (Negative Results)**

The authors tested several potential fixes to close this gap:

  • In-Model Fix: They introduced a discrete-aware cross-coupling head to condition categorical predictions on the other columns. This intervention produces no significant change on any of the three datasets we test, with the standard deviation being at least as large as the mean, indicating a failure of this remedy.

  • Higher-Order Structure: The residual gap is determined to be higher-order. A second-moment or covariance-matching fix was deemed inert because the mismatch lives in higher-order structure that such a term does not see.

** Conclusion and Takeaway**

The paper concludes that the gap is not closed by any of the tested interventions:

  • None of the cheap interventions we test closes the gap.

  • The residual gap is higher-order, so a covariance-matching fix has nothing to correct.

  • The diagnostic tool shows that while dependency is necessary for utility (the FF collapse), the generators' shortfall does not track this measured gap.

Guidance for Practitioners

The authors advise practitioners to:

  1. Report the dependency decomposition with fully-factorized and oracle references, and do not rely on Trend or the linear (logistic-regression) C2ST to certify dependency.

  2. Be cautious against assuming that an added module or more scale will help, as both a 16 times capacity increase and a discrete-aware cross-coupling module failed to close the gap.

The ultimate suggestion for method designers is that, given the evidence, supervising dependency directly in the objective is a natural next intervention to test.

Improvements for AI systems

Based on my analysis of this paper, I have identified several critical areas where current AI systems are failing in tabular data generation and provide specific, actionable improvements.

The core problem identified is that existing fidelity metrics (like linear C2ST and Pairwise Trend scores) are fundamentally blind to inter-column dependency. This means a synthetic dataset can be judged nearly real by standard benchmarks even if its crucial joint structure is completely destroyed.

The improvements below focus on integrating the authors' diagnostic framework into both system evaluation and generative model design.


We must replace reliance on single, aggregate metrics with a decompositional diagnostic that localizes structural deficits within a single score.

Specific Improvements:

  • Adopt the XGBC2ST Framework: Implement the gradient-boosted tree discriminator (XGB-C2ST) as the primary detection mechanism, rather than a simple linear classifier.

  • Mandatory Decomposition: The final fidelity score must be decomposed into three distinct, observable components:

  1. Marginal Fidelity (marg): Measures discriminability due to mismatched column marginal distributions (the easy part).

  2. Dependency Gap (dep): Measures additional discriminability arising from faulty joint structure (the hard part). This is the primary indicator of failure.

  3. Cross-Coupling Fidelity (dep cross): Measures how well the system preserves the interaction between numerical and categorical features, isolated by blocking out within-block dependencies.

What the Improved System Can Do:

  • Pinpoint Failure Modes: The system can instantly distinguish between a generator that has successfully matched marginal distributions (high marg) but failed to capture joint structure (low dep), providing clear, actionable feedback to engineers.

  • Quantify Blind Failures: It will flag generators that the linear C2ST would rate as passing, immediately highlighting their lack of structural fidelity.

The paper demonstrates that current generative objectives (e.g., moment-matching in EF-VFM, additive denoising in TabDiff) only supervise per-dimension posterior means, not the joint structure—this is a fundamental design flaw.

The paper shows that simple fixes (like post-hoc copula correction or brute-force scaling) are insufficient because the gap is often higher-order and not a function of capacity alone.

Feature Old System (Current) New Improved System Key Advantage

:---:---:---:---

Fidelity Check C2ST/Trend (Single, aggregate score) dep and dep cross Decomposition (Diagnostic) Pinpoints where the data fails structurally.

Design Goal Match marginal statistics (Mean-field objective) Supervise Joint Cumulants (Joint dependency term) Ensures utility for complex tasks like minority-class detection.

Failure Response Assume failure is due to lack of capacity or poor training. Detect Higher-Order structural deficit. Provides a precise, non-obvious diagnosis and prevents wasted effort on insufficient remedies.

Sources

Related papers