FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching".
Jane: The paper was written by Suman Cha, Seongchan Lee, Dohyun Ko and Hyunjoong Kim from Yonsei University and KAIST.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper summary: Tom: Today's paper comes from Yonsei University and KAIST, and it's called FUSE — Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching. The authors are Suman Cha, Seongchan Lee, Dohyun Ko, and Hyunjoong Kim. In one sentence, they've built a new neural architecture for generating synthetic tabular data, and it posts consistently strong results across eight benchmark datasets.
Jane: What grabs me is the split in the architecture before we even look at those numbers. One module specializes the computation per feature type, and a separate mechanism handles exchange across columns.
Tom: That split is the whole thesis. And the results back it up — best average rank on marginal fidelity at 1 point 25, and best rank on the multivariate classifier test at 1 point 62.
Lu: Synthetic tabular data matters because hospitals and banks use it to share records without exposing real people, and it can help when your training set is thin. Data sharing, scarcity, and robust downstream analysis are the motivations the paper lists up front.
Meng: The hard part is that tables mix very different kinds of variables. A column for income behaves nothing like a column for employment status, and even two numerical columns can look completely different — one heavily skewed, another multimodal.
Jane: The paper builds on variational flow matching, where each column gets its own endpoint distribution. Numerical columns get Gaussian factors, categorical columns get categorical ones. But the network behind it all is a shared backbone, and feature identity is just an embedding.
Tom: So the loss treats columns individually, but the computation doesn't.
Lalam: That's the gap FUSE closes. Type-specific adaptive mixture modules let each feature combine shared subnetworks in its own way, and joint self-attention lets numerical and categorical columns exchange information across the entire row. Specialization and exchange become explicit, separate mechanisms.
Tom: There's a theoretical side as well. They show what you pay when a predictor can't condition on the full context, and they bound the Wasserstein distance between generated and real data in terms of endpoint prediction risk.
Lalam: An architectural idea, supporting theory, and consistent experimental gains — that combination is why this paper deserves a careful read.
Jane: The introduction on page one explains why mixed-type generation is still hard, so let's start there.
Page 1: Tom: We've heard the one-sentence version of the paper. The introduction on page one lays out exactly why mixed-type generation remains hard.
Jane: Page one opens with the motivation — tabular data dominates machine learning applications in healthcare, finance, and public policy, and synthetic copies make data sharing and augmentation possible without exposing real records. The paper also brings up mitigating data scarcity and enabling robust downstream analysis.
Lu: Then comes a point that shapes the whole paper. The difficulty isn't just numerical versus categorical. Columns within the same type have different structures — a numerical column can be skewed, heavy-tailed, or multimodal, and categorical columns differ in cardinality and in how concentrated their frequencies are.
Tom: So even two columns of the same type shouldn't necessarily get the same treatment.
Meng: The paper walks through the flow matching lineage. Flow matching itself learns a time-dependent velocity field, and integrating that field transports a simple source distribution to the data distribution.
Jane: Variational flow matching expresses that velocity through inference over endpoints, and the mean-field version factorizes that inference across columns. Then exponential-family VFM assigns appropriate distributions to each variable type.
Lu: Each step makes the objective more column-aware, which is exactly why the architectural gap becomes visible. The factorization aligns the objective with the data types, but nothing in the network does.
Tom: Which brings us to TabbyFlow, the direct predecessor. It implements EF-VFM with a shared backbone, processing column embeddings through a common network that doesn't distinguish feature types in its computation.
Jane: The authors' argument is that column embeddings encode identity, but they don't provide any explicit mechanism for feature-dependent transformations. That's the gap.
Lalam: So the motivation is architectural. The loss already respects column types, but the network doesn't. And the fix has to preserve parameter sharing while adding specialization — which is exactly the balance FUSE strikes.
Tom: The introduction closes with three contributions — the FUSE architecture, theory around restricted conditioning and Wasserstein generation error, and the experimental evaluation.
Jane: Before we can judge those contributions, page two builds the formal framework underneath.
Page 2: Tom: Page two supplies the formal machinery. The endpoint X1 stacks the numerical columns together with one-hot vectors for the categorical columns, so the overall space has dimension equal to the numerical count plus all the category counts.
Jane: The source distribution is standard Gaussian on that space, and the conditional path is linear — Xt is one minus t times the noise plus t times the data point. The conditional velocity is simply the endpoint minus the state, divided by one minus t.
Lu: The clean consequence follows immediately. The marginal velocity only depends on the endpoint posterior through its mean, so everything else about the posterior is irrelevant to the flow.
Meng: That's what makes variational flow matching tractable. Instead of modeling the full joint posterior over endpoints, you approximate it, and the mean-field version factorizes that approximation across columns.
Tom: And exponential-family VFM picks the right family per column — Gaussian factors for numerical endpoints, categorical factors for categorical ones, with a predetermined variance schedule.
Jane: So what does the training objective look like? It's a mean-squared error for the numerical columns plus a cross-entropy term for each categorical column. The categorical term is the log-probability of the true category under the predicted distribution.
Tom: Here's the subtle part. The mean-field assumption factorizes the variational distribution across columns, but it places no restriction on the conditioning argument. Each predictor can still see every coordinate of the intermediate state.
Lalam: That distinction matters. The factorization shapes the output distribution, but the network's capacity to combine information is left completely open.
Meng: And the paper says it plainly — the variational objective does not specify how computation should be shared among these column-wise parameter functions. That's the slot FUSE fills.
Jane: Yes. Page three shows the architecture that fills that slot.
Page 3: Tom: Page three introduces the architecture, and the first thing to notice is that every feature becomes a token. A feature-specific linear projection, plus a time embedding, tells the network where it is along the flow.
Jane: The adaptive mixture module is the centerpiece. It computes alignment scores between feature tokens and a set of latent components, then normalizes those scores along two axes.
Lu: Two normalizations, two different jobs. Aggregation weights, normalized across features, decide how tokens get pooled into components. Recombination weights, normalized across components, decide how each feature pulls information back out.
Meng: So a feature like age sends a weighted mixture of its own representation into a handful of shared subnetworks, and the output comes back recombined according to a different set of weights. Different columns end up using the subnetworks in different proportions.
Tom: And the subnetworks are shared across all features of the same type, which keeps the parameter count under control. You get specialization without a separate network for each column.
Jane: The numerical and categorical types never share subnetworks with each other, which is where the type-specificity comes from.
Lalam: Then joint attention takes the processed numerical and categorical tokens, concatenates them, and applies standard multi-head self-attention. That's the exchange mechanism — every column can condition on every other column.
Tom: The two roles are now explicit. Mixtures for specialization, attention for exchange.
Jane: Endpoint heads read out the numerical means and categorical probabilities from the final tokens, and those predictions define the velocity field.
Meng: One detail I appreciate is that the normalization in each branch is time-conditioned, so the balance between components can vary along the flow. The same architecture can behave differently early and late in generation.
Tom: The architecture is in place. Page four asks what goes wrong if you remove either piece.
Page 4: Tom: Page four is theory. Proposition one quantifies the cost of restricted conditioning — if a predictor can't see some features, the excess risk decomposes into a numerical term plus a sum of KL divergences.
Jane: The numerical term is the expected squared gap between the full and restricted posterior means, scaled by the noise variance. Each categorical term measures how much posterior information about that category is lost.
Lu: Then come two bottleneck examples that separate two failure modes.
Tom: The conditioning bottleneck shows that an informative numerical proxy isn't sufficient. If an endpoint's mean depends on a latent sign, and the predictor only sees a noisy version of that sign, the excess risk is strictly positive — they compute it as a constant times the expected variance of the sign given the proxy.
Meng: The representation bottleneck is about shared computation. Two endpoint means that live in orthogonal directions can't be captured by a single shared scalar representation; the minimum error is the smaller of the two squared coefficients.
Jane: But a shared dictionary with two components represents both means exactly, as long as features use different recombination weights. The paper gives explicit weights — three quarters and one quarter, swapped between the two features.
Lalam: So the two components of FUSE map onto two distinct failure modes. Attention addresses the conditioning bottleneck, and mixtures address the representation bottleneck.
Tom: Theorem one then connects endpoint risk to generation quality. Under Lipschitz and integrability assumptions, the Wasserstein error is bounded by the square root of the excess risk, scaled by a constant that depends on the time horizon, plus a truncation term.
Jane: The trade-off is explicit. Larger T shrinks the truncation term because the flow gets closer to the final time, but it inflates the coefficient in front of the risk term.
Lu: The appendix verifies that a fixed FUSE network satisfies the regularity conditions — the vector field is bounded, Lipschitz, and integrable at the origin. So the bound applies to their actual model.
Tom: With the theory in hand, the paper moves to experiments. Page five sets up the evaluation.
Page 5: Tom: Page five sets up the experimental campaign. Eight datasets — Adult, Default, Beijing, Shoppers, Magic, News, Diabetes, and Fault — spanning binary classification, multiclass classification, and regression.
Jane: Sample sizes range from 691 to 37,581. Seven baselines across four families: CTGAN for GANs, TVAE for VAEs, CoDi, TabDDPM, TabSyn, and TabDiff for diffusion, and TabbyFlow for flow matching.
Lu: Four metrics cover different quality dimensions. Shape measures marginal fidelity with Kolmogorov-Smirnov statistics for numerical features and total variation distances for categoricals.
Meng: Trend measures pairwise dependencies by comparing correlation structures. C2ST trains an XGBoost classifier to tell real from synthetic, and MLE trains a model on synthetic data and scores it on held-out real data.
Tom: The headline is Shape, where FUSE ranks first on six datasets and second on the other two — an average rank of 1 point 25, against 2 point 50 for TabbyFlow.
Jane: That gap is striking, and it's consistent. The improvement appears on every dataset, not just one or two favorable ones.
Lalam: Shape is also the metric most directly tied to their design. If type-specific processing works, it should show up first in the marginals.
Tom: The main tables average over twenty random seeds, and the implementation uses a hidden dimension of 256 with four layers and four attention heads, trained with AdamW. So these aren't cherry-picked runs.
Jane: Page six asks whether the advantage survives when the metrics look at dependencies and downstream tasks.
Page 6: Tom: Page six reports Trend, C2ST, and MLE. On Trend, FUSE has the highest score on five datasets and the best cross-dataset mean, 0 point 982 against 0 point 978 for TabbyFlow.
Jane: The largest Trend gains over TabbyFlow are on Diabetes and Fault — from 0 point 936 up to 0 point 966, and from 0 point 978 up to 0 point 985.
Lu: But TabbyFlow still holds a slightly better average rank on Trend, 1 point 75 to 2 point 00, because it ranks higher on Beijing, Magic, and News. So the pairwise dependence story is competitive rather than dominant.
Tom: C2ST is where FUSE stands out. Average rank 1 point 62, first on Default, Beijing, and News, second on the remaining five. It's the only method that places in the top two on all eight datasets.
Meng: That consistency matters because C2ST uses all features jointly. If FUSE only nailed univariate marginals, a classifier looking at the full row would still find discrepancies.
Jane: On MLE, FUSE has the best average rank at 2 point 62, and the standout result is News — RMSE drops from 0 point 872 for TabbyFlow to 0 point 841. That's roughly a 3 point 6 percent relative improvement.
Tom: The visualizations back up the tables. Correlation error heatmaps show FUSE at near-zero errors on News, while other methods show widespread discrepancies.
Lu: There are wins on Shoppers and Fault too, plus a second-best on Default — the utility gains aren't concentrated in one dataset.
Tom: So the aggregate picture is strong. The next page shows which components drive it.
Page 7: Tom: Page seven contains the component analysis. Six configurations vary whether the mixture modules are active, whether attention is joint or restricted, and whether mixtures apply to numerical, categorical, or both types.
Jane: The joint attention effect is dramatic under dense processing. Enabling it lifts Trend from 0 point 960 to 0 point 976, and AUC from 0 point 630 to 0 point 888 — that's a massive jump in downstream utility.
Lu: With both mixture modules active, the comparison is similar. Trend goes from 0 point 960 to 0 point 977, AUC from 0 point 644 to 0 point 882, and RMSE from 0 point 818 down to 0 point 725.
Meng: The adaptive mixture modules contribute elsewhere. Under joint attention, they push Shape from 0 point 984 to 0 point 985, C2ST from 0 point 985 to 0 point 991, and β-Recall from 0 point 482 to 0 point 545.
Tom: So the picture is clear. Attention is the workhorse for dependencies and downstream utility, while the mixtures improve multivariate fidelity and coverage.
Jane: The paper also notes trade-offs. The numerical-only mixture achieves the highest β-Recall, and the categorical-only version matches FUSE on RMSE, but the full configuration has the strongest overall profile.
Lalam: That division of labor was predicted by page four's theory. Attention prevents the conditioning bottleneck, mixtures prevent the representation bottleneck.
Tom: The conclusion draws it together — FUSE makes the parameter-sharing configuration explicit, the theory characterizes the roles of each component, and the experiments confirm they're complementary.
Jane: The supplementary material extends this with a controlled experiment where they dial the cross-type dependence from zero to nearly one. The advantage of joint attention grows right along with it.
Lu: For the field, the larger message is that tabular generation benefits from an architecture that matches the data structure.
Tom: Let's close with what this means beyond the benchmark tables.
Conclusion: Tom: So where does this leave us? FUSE separates feature specialization from cross-column exchange in variational flow matching, and it does so with a clean modular design that's easy to describe.
Jane: The theory gives a vocabulary for why it works. Restricted conditioning costs measurable risk, and endpoint prediction quality directly bounds generation quality. That's a useful bridge between the loss you optimize and the distribution you actually generate.
Lu: The empirical record is strong — best average rank on Shape and C2ST, competitive on Trend and MLE. And the consistency across eight datasets is the most convincing part of it.
Meng: The ablations convince me most. Each component earns its place, and the theory anticipated which metric each component would affect — attention for dependencies, mixtures for fidelity and coverage.
Lalam: For the wider field, this is another sign that tabular generation is moving from generic backbones toward purpose-built designs. When you know the data structure, you should exploit it, and here the structure is mixed types.
Tom: The practical implications reach beyond benchmarks. Better synthetic tables mean safer data sharing in healthcare and finance, and more reliable augmentation when real data is scarce.
Jane: There's room to push further. The theory applies before categorical decoding, and the controlled dependence experiment in the appendix points at a deeper analysis of when attention pays off.
Lu: The paper's own framing is apt — a favorable balance across marginal, pairwise, multivariate, and task-oriented evaluations, without relying on dominance in any single metric.
Meng: Variational flow matching is still young, and this paper shows the objective and the architecture can be developed almost independently. That separation is a useful blueprint.
Tom: For anyone working with synthetic data, this paper is worth a careful read. We'll be watching for follow-ups.
Jane: That closes our discussion. On to the next one.
Suman Cha, Seongchan Lee, Dohyun Ko, Hyunjoong Kim
Yonsei University · KAIST
cs.LG, cs.AI
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 19 pages, 7 figures, 7 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 67/100
The gist: The paper addresses the challenge of generating synthetic mixed-type tabular data, which "requires capturing the complex joint distributions of mixed-type variables." The authors note that "the
Key concepts
- Flow matching
- A generative modeling technique that learns a velocity field to transport a simple distribution (like Gaussian noise) to the data distribution. The field is integrated over time to produce new samples. Variational flow matching approximates the posterior over endpoints to make this tractable.
- Mixed-type tabular data
- Data tables that contain both numerical and categorical columns, each with different distributions and structures. Generating synthetic versions of such data is challenging because a single model must handle diverse feature types and their dependencies.
- Adaptive mixture module
- A component in FUSE that lets each feature combine shared subnetworks using learned weights. It normalizes scores across features and components to pool and recombine information, allowing specialization without separate networks per column.
- Cross-column exchange
- A mechanism in FUSE using joint self-attention over all feature tokens, enabling each column to condition on every other column in the row. This facilitates capturing dependencies between numerical and categorical features.
Terminology
Summary
The paper addresses the challenge of generating synthetic mixed-type tabular data, which requires capturing the complex joint distributions of mixed-type variables.
The authors note that "the challenge extends beyond the distinction between numerical and categorical variables. Columns within the same type can also exhibit different statistical structures. Numerical columns may vary in skewness, tail behavior, and multimodality, while categorical columns differ in cardinality and frequency concentration."
The work builds on Variational Flow Matching (VFM) and its extensions. "Variational Flow Matching (VFM) expresses the velocity through conditional inference over trajectory endpoints, while Mean-Field Variational Flow Matching (MF-VFM) makes the endpoint inference tractable by factorizing the variational endpoint distribution across variables. For mixed-type data,
Exponential Family Variational Flow Matching (EF-VFM) extends this approach by assigning appropriate exponential family distributions to match the statistical properties of each variable type."
The authors identify a key gap: "While endpoint factorization aligns the objective with varying data types, the architectural challenge of processing diverse variables remains unresolved. TabbyFlow implements EF-VFM by processing column embeddings through a shared backbone. Although this design supports parameter sharing and encodes feature identity through column embeddings, it provides no explicit mechanism for feature-dependent transformations."
The paper introduces FUSE (Feature-wise Unified Specialization with cross-column Exchange), an architecture interleaving type-specific adaptive mixture processing with joint attention, enabling feature specialization while maintaining cross-column dependencies.
Within each variable type, adaptive mixture processing forms feature-dependent combinations of shared specialized subnetworks using a differentiable aggregate-transform-recombine operator.
At each layer, alignment scores are computed, then normalized along two axes
: aggregation weights (normalized across feature tokens for each component) and recombination weights (normalized across components for each feature). The aggregation weights form latent components, each of which is transformed by a shared specialized subnetwork, and the recombination weights distribute these transformed components back to the feature tokens, producing the residual update.
Since both weights depend on the feature representations, each feature can receive a distinct combination of subnetwork transformations,
while the subnetworks are also shared across all features of type r, allowing type-specific specialization without introducing a separate subnetwork for every feature.
After adaptive mixture processing, we concatenate the type-specific outputs and apply multi-head self-attention,
so that each feature token can incorporate information from all numerical and categorical features.
This allows cross-type information exchange and unrestricted conditioning for every endpoint predictor.
After L blocks, feature-specific output heads map the final numerical and categorical tokens to the parameters of their corresponding endpoint factors
—Gaussian factors for numerical endpoints and categorical factors for categorical endpoints, optimized using the endpoint objective.
The paper provides three theoretical contributions:
1. Conditioning Penalty (Proposition 1): The paper quantifies the exact excess population risk induced by restricting the conditioning context. For a restricted conditioning σ-algebra G versus full conditioning F:
"The decomposition vanishes when the restricted context preserves both the numerical posterior mean and every categorical posterior probability. A penalty arises only when the omitted features contain residual endpoint information after conditioning on the retained features."
2. Architectural Bottleneck Analysis: The paper distinguishes two failure modes via illustrative constructions:
-
Conditioning bottleneck: "an informative numerical proxy for S need not be sufficient for the endpoint mean when σ2 > 0. Joint attention allows the categorical representation to enter the numerical feature update and therefore avoids imposing the restricted conditioning."
-
Representation bottleneck: A single shared scalar representation
restricts both endpoint functions
and incurs positive approximation error, whereasfeature-dependent recombination can remove the approximation error induced by the specified single-representation restriction while retaining a shared set of transformations.
3. Distributional Error Bound (Theorem 1): The following theorem bounds the Wasserstein error of the generated distribution in terms of ET (f)
(the excess endpoint risk):
"The first term converts endpoint excess risk into a Wasserstein error bound, while the second captures truncation at T < 1. Hence, larger T reduces truncation error but increases the risk-dependent coefficient."
The Appendix verifies that a fixed FUSE network satisfies the regularity conditions used in Theorem 1,
including uniform Lipschitz continuity of the vector field.
The method is evaluated on "eight tabular datasets: Adult, Default, Beijing, Shoppers, Magic, News, Diabetes, and Fault. Each dataset contains both numerical and categorical features and is associated with either binary classification, multiclass classification, or regression. It is compared against
seven competitive synthetic tabular data generation methods spanning four model families: GAN-based (CTGAN), VAE-based (TVAE), Diffusion-based (CoDi, TabDDPM, TabSyn, TabDiff), and Flow-based (TabbyFlow)."
Metrics include Shape (marginal distributional similarity), Trend (pairwise feature dependencies), C2ST (Classifier Two-sample Test), Machine Learning Efficiency (MLE), plus α-Precision and β-Recall in the appendix.
FUSE demonstrates strong and consistent performance:
-
Shape:
It ranks first on six datasets and second on the remaining two in Shape, resulting in an average rank of 1.25. The next-best method, TabbyFlow, obtains an average rank of 2.50.
-
Trend:
FUSE achieves the highest score on five datasets and the highest cross-dataset mean of 0.982, compared with 0.978 for TabbyFlow.
-
C2ST:
FUSE achieves the best C2ST average rank of 1.62, followed by TabbyFlow at 1.75... it is the only method that places among the top two across all eight datasets.
-
MLE:
FUSE also achieves the best average MLE rank of 2.62, outperforming TabbyFlow at 2.75 and TabDiff at 3.12.
On the News dataset, itreduces RMSE from 0.872 for TabbyFlow... to 0.841. This corresponds to a relative reduction of approximately 3.6%.
Marginal distribution plots show that FUSE recovers both the global shape and local structure of the empirical distributions,
including the sequence of local modes in the polarity variable on News and Shoppers,
and for categorical features, FUSE closely matches the empirical category frequencies across Adult, Beijing, Magic, and News.
Correlation heatmaps show FUSE achieves consistently low correlation errors across all blocks, with a pronounced advantage on News and Fault.
Ablation studies evaluate six configurations. The results show that joint attention enables effective cross-type conditioning, while adaptive mixture processing further improves fidelity and coverage through type-specific computation.
Specifically, under dense processing, enabling joint attention preserves Shape at 0.984 while increasing Trend from 0.960 to 0.976 and AUC from 0.630 to 0.888,
consistent with Proposition 1. Adaptive mixture processing under joint attention improves Shape from 0.984 to 0.985, Trend from 0.976 to 0.977, C2ST from 0.985 to 0.991, and β-Recall from 0.482 to 0.545.
A controlled cross-type dependence experiment (in the appendix) further shows that both the population penalty and the learned endpoint-risk reduction are near zero at independence and increase with cross-type dependence,
and the MLE gain increases from 0.003 to 0.508 at ρdep = 0.95.
Additionally, the relative increase in categorical endpoint risk under restricted attention
is positive on every dataset across all seeds, confirming that joint attention becomes most consequential for downstream utility as task-relevant cross-type dependence strengthens.
The authors conclude: "FUSE combines type-specific adaptive mixture processing with joint attention. The adaptive mixture modules form feature-dependent combinations of shared specialized subnetworks, while joint attention preserves cross-type information exchange and unrestricted conditioning for every endpoint predictor. Our theoretical analysis characterizes the complementary roles of these components... Comprehensive experiments show that FUSE achieves strong aggregate performance across the evaluated metrics."
Improvements for AI systems
- Mixed-type tabular data generators with feature-specialized, parameter-efficient architectures.
Instead of forcing all columns through the same backbone or assigning a separate network to each column, an improved system can use adaptive mixture processing: each type (numerical, categorical) shares a small set of specialized subnetworks, and each feature learns its own differentiable combination of those subnetworks. The result is a generator that captures column-specific skewness, multimodality, tail behavior, cardinality, and frequency concentration without exploding parameter count.
- Generative models with explicit cross-type conditioning.
Add joint multi-head self-attention over all feature tokens after type-specific processing. This gives every endpoint predictor unrestricted access to both numerical and categorical features, avoiding the conditioning bottleneck where informative features cannot influence endpoint predictions for other variables. The improved system can model cross-type dependencies (e.g., a categorical category affecting a numerical distribution's mean or variance) much more faithfully.
- Principled endpoint optimization with a tunable time horizon.
Train using the variational flow-matching endpoint objective, with the flow time horizon T treated as a controllable trade-off: larger T lowers truncation error, smaller T lowers the risk-dependent Wasserstein coefficient. An improved system can tune T per dataset to balance generation fidelity and endpoint risk, and it can monitor the theoretical distributional error bound as a training diagnostic.
- Feature-dependent recombination instead of homogeneous residual updates.
Replace fixed residual blocks with an aggregate-transform-recombine operator where aggregation weights normalized across features and recombination weights normalized across components allow each feature to receive a distinct transformation while still sharing computation. This removes representation bottlenecks that occur when a single shared representation must serve multiple conflicting endpoint functions.
- Better synthetic tabular data across all evaluation axes.
The improved system can generate synthetic mixed-type data that:
-
matches marginal distributions better (lower Shape error, rank 1.25 vs. next-best 2.50),
-
preserves pairwise feature dependencies more accurately (Trend 0.982 cross-dataset mean),
-
is significantly harder to distinguish from real data in a classifier two-sample test (best average C2ST rank, always top two across datasets),
-
yields better downstream machine-learning efficiency, e.g., reducing RMSE on the News regression task from 0.872 to 0.841 (3.6% relative improvement).
- Controlled handling of cross-type dependence strength.
With the theoretical conditioning-penalty result, the improved system can quantify when joint attention matters most: the penalty is zero when retained features preserve all relevant posterior information, and grows with true cross-type dependence. This allows an adaptive system to selectively enable or emphasize joint attention when task-relevant cross-type dependence is high, improving downstream utility (MLE gain from 0.003 at independence to 0.508 at strong dependence).
- Generative tabular models with verifiable regularity conditions.
The architectural choices — uniform Lipschitz vector fields, bounded components, fixed network parameters — satisfy the regularity conditions required for the Wasserstein error bound. An improved system can therefore provide formal guarantees linking endpoint excess risk to generated-distribution quality, rather than relying on heuristic evaluation alone.
Abstract
Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Variational flow matching handles distinct endpoints via factorized distributions, yet leaves feature-specific processing and cross-column interactions implicit within a shared backbone. We introduce Feature-wise Unified Specialization with cross-column Exchange (FUSE) to explicitly separate these roles. FUSE applies separate adaptive mixture modules to numerical and categorical features, allowing each feature to combine shared specialized subnetworks, while joint attention preserves information exchange across all columns. We also characterize the excess population risk from restricted conditioning contexts and bound the continuous Wasserstein generation error by endpoint-prediction risk. Comprehensive experiments on eight tabular datasets demonstrate that FUSE achieves strong and consistent performance across distributional fidelity and downstream utility metrics.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks