Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations

arXiv:2511.04000 · cs.LG, cs.AI, cs.CL, stat.ML · Submitted 2025-11-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations".

Jane: The gist The work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees.

Tom: First, who's behind it and why it matters.

Paper summary: Jane: So we've covered the basic thesis: using synthetic data generation to train a model to create good decision trees efficiently. Now let's dig into the actual process they laid out in this paper, "Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations."

Tom: Okay, so the paper outlines a four-step pipeline for generating both the synthetic data and those corresponding near-optimal trees simultaneously >

Lu: It starts with sampling synthetic features and target labels from a Structural Causal Model, which is what ensures those causal relationships are there from the beginning >

Meng: Then they generate CART trees using these synthetic datasets to set a baseline for performance, but they immediately apply quality filters in step three to toss out anything with severe class imbalance or poor separability >

Jane: That filtering stage is really important because it ensures the data going into the next step is actually usable for training decision trees, not just noise >

Lalam: And finally, in step four, they create the actual synthetic datasets that line up perfectly with those filtered decision boundaries by relabeling and adding some controlled label noise >

Tom: The core mechanism here is using the MetaTree transformer architecture to predict near-optimal decision trees based on these synthetic inputs during the meta-learning step >

Jane: So, they are using this structure to learn how to generate those good trees without having to solve every single problem from scratch on real data >

Conclusion: Tom: So wrapping up, the title "Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations" points to a focus on making decision trees scalable through this synthetic generation technique.

Jane: It suggests they're moving away from the traditional bottleneck of needing massive amounts of real data or complex optimal solvers for every single problem >

Lu: The implication is that we can generate training targets that are tailored to decision tree construction right from the start, which should make meta-learning these models much more efficient >

Meng: From an engineering standpoint, it means less reliance on computationally heavy tree solvers and more use of this synthetic data pipeline for pre-training >

Lalam: And for us in the AI space, this gives us a way to generate high-quality training targets systematically without needing to curate huge datasets manually or rely on extremely expensive real-world optimal trees >

Tom: So, essentially, they're showing a path toward building interpretable decision tree models that are both scalable and efficient by learning from synthetic data generated via an SCM workflow >

Capital One

cs.LG, cs.AI, cs.CL, stat.ML

Submitted: 2025-11-06

Updated: 2026-10-07

Comments: 9 pages, 3 figures, Neurips 2025 GenAI in Finance Workshop

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The gist The work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees.

Key concepts

Meta-learning Workflow
This involves two main steps: first, meta-learning where the model learns by training on labeled synthetic datasets and their optimal decision trees. Second, an inference step where the pre-trained model predicts near-optimal trees for new, unseen real-world data.
Structural Causal Model (SCM)
An SCM is used to generate synthetic data by defining causal relationships between features and labels. This ensures that the generated datasets maintain realistic causal connections, which is crucial for training decision trees effectively.
Quality Filters
These filters select high-quality synthetic datasets for pre-training. They include a class imbalance filter (ensuring no majority class exceeds 75%) and an accuracy filter (retaining only datasets where a CART tree can achieve over 70% accuracy). These filters favor smaller numbers of classes.
Synthetic Data Generation Pipeline
This four-step process creates synthetic data: sampling from SCMs, generating baseline CART trees, applying quality filters, and finally relabeling the data while introducing 5% label noise. This process aligns the synthetic datasets with decision boundaries for effective pre-training.

Terminology

Summary

The gist The work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees.

Meta-learning Workflow

  1. meta-learning (or pre-training) step where labeled synthetic datasets are fed into MetaTree along with the optimal decision trees for each dataset as the training targets.

  2. inference step where the pre-trained model is used to predict the near-optimal, look-ahead trees on an unseen, real-world datasets.

The MetaTree model is pre-trained on synthetic datasets generated using a Structural Causal Model (SCM) workflow. The SCM ensures causal relationships between features and labels, while the pipeline filters out low-quality datasets using class imbalance and accuracy thresholds.

Synthetic Data Generation

The four-step pipeline for generating synthetic data and corresponding near-optimal trees is outlined as follows:

(1) Structural Causal (SC) Graphs.

(2) CART Decision Boundaries.

(3) Apply Quality Filters.

(4) Label Assignment and Noising.

This process involves sampling synthetic features and target labels from a Structural Causal Model (SCM) to ensure causal relationships between features and labels. CART trees f(X) are generated using the synthetic datasets to establish a baseline for decision tree performance. Quality filters are applied to remove datasets with issues like severe class imbalance or poor separability. Finally, synthetic datasets (X, y') are created aligned with the decision boundaries of the CART trees by relabeling original labels and introducing 5% label noise.

Quality Filters

To enhance raw synthetic data for decision tree construction, two quality filters are introduced:

** A class imbalance filter ensures that the majority class comprises no more than 75% of the samples by filtering for SCM datasets with an Inormalized value less than 0.3. This metric is defined using the Penn Machine Learning Benchmarks (PMLB) [Olson et al. (2017)]. The filter prevents trivial decision stumps by ensuring no majority class exceeds 75% of the samples. A second accuracy filter retains only datasets where a CART tree could achieve over 70% accuracy. These measures select suitable datasets for generating effective pre-training targets for meta-learning. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmarks. Furthermore, the class count distribution shows a systematic decline as the number of classes increases. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmark. The class count distribution shows a systematic decline as the number of classes increases. Both accuracy and class imbalance filters will favor datasets with smaller number of classes. This trend is an intended consequence of our accuracy and class imbalance quality filters, which inherently favor datasets with fewer classes. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmark. Furthermore, the class count distribution shows a systematic decline as the number of classes increases. Both accuracy and class imbalance filters will favor datasets with smaller number of classes. This trend is an intended consequence of our accuracy and class imbalance quality filters, which inherently favor datasets with fewer classes. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmark. Furthermore, the class count distribution shows a systematic decline as the number of classes increases. Both accuracy and class imbalance filters will favor datasets with smaller number of classes. This trend is an intended consequence of our accuracy and class imbalance quality filters, which inherently favor datasets with fewer classes. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmark. Furthermore, the class count distribution shows a systematic decline as the number of classes increases. Both accuracy and class imbalance filters will favor datasets with smaller number of classes. This trend is an intended consequence of our accuracy and class imbalance quality filters, which inherently favor datasets with fewer classes. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmark. Furthermore, the class count distribution shows a systematic decline as the number of classes increases. Both accuracy and class imbalance filters will favor datasets with smaller number of classes. This trend is an intended consequence of our accuracy and class imbalance quality filters, which inherently favor datasets with fewer classes. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmark. Furthermore, the class count distribution shows a systematic decline as the number of classes increases. Both accuracy and class imbalance filters will favor datasets with smaller number of classes. This trend is an intended consequence of our accuracy and class imbalance quality filters, which inherently favor datasets with fewer classes. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmark. Furthermore, the class count distribution shows a systematic decline as the number of classes increases. Both accuracy and class imbalance filters will favor datasets with smaller number of classes. This trend is an intended consequence of our accuracy and class imbalance quality filters, which inherently favor datasets with fewer classes. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmark. Furthermore, the class count distribution shows a systematic decline as the number of classes increases. Both accuracy and class imbalance filters will favor datasets with smaller number of classes. This trend is an intended consequence of our accuracy and class imbalance quality filters, which inherently favor datasets with fewer classes.

Improvements for AI systems

  1. The system can generate large-scale, realistic synthetic datasets to enable meta-learning of decision trees, as demonstrated by Our approach samples near-optimal decision trees synthetically, creating large-scale, realistic datasets. This allows for pre-training models like MetaTree on diverse data without relying on scarce real-world data.

  2. The system can circumvent the computational bottlenecks of traditional optimal tree solvers by using a constructive method that generates synthetic data and corresponding near-optimal trees simultaneously. This is achieved through a pipeline that uses a Structural Causal Model (SCM) to sample features and labels, followed by CART tree generation and subsequent label reassignment.

  3. The system can achieve performance comparable to pre-training on real-world data or with computationally expensive optimal decision trees by utilizing the MetaTree transformer architecture trained on the synthetic data, as shown in Figure 1(1).

  4. The system can ensure that generated training targets are suitable for decision tree construction by implementing quality filters, specifically ensuring no majority class exceeds 75% of the samples via a normalized class imbalance metric and retaining only datasets where a CART tree could achieve over 70% accuracy.

Abstract

Decision trees are widely used in high-stakes fields like finance and healthcare due to their interpretability. This work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees. Our approach samples near-optimal decision trees synthetically, creating large-scale, realistic datasets. Using the MetaTree transformer architecture, we demonstrate that this method achieves performance comparable to pre-training on real-world data or with computationally expensive optimal decision trees. This strategy significantly reduces computational costs, enhances data generation flexibility, and paves the way for scalable and efficient meta-learning of interpretable decision tree models.

Sources

Related papers