From Uniform to Learned Knots: A Study of Spline-Based Numerical Encodings for Tabular Deep Learning

arXiv:2604.05635 · cs.LG · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Uniform to Learned Knots: A Study of Spline-Based Numerical Encodings for Tabular Deep Learning".

Jane: The paper was written by Manish Kumar, Anton Frederik Thielmann, Christoph Weisser and Benjamin Säfken from BASF and Clausthal University of Technology and Amazon Music and Bielefeld School of Business, Hochschule Bielefeld.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everyone! Today we're digging into a paper with a title that's a mouthful: "From Uniform to Learned Knots: A Study of Spline-Based Numerical Encodings for Tabular Deep Learning." Jane, when you first saw that title, what jumped out at you?

Jane: Tom, it was the word "knots" that got me. In math, a knot isn't about tying rope. It's a point where a curve changes its shape, like where a piecewise function bends. And the paper is asking whether we should just pick those bend points randomly, or let the model learn where they should be.

Lu: Exactly, Jane. And that's a bigger deal than it sounds. Tabular data—think spreadsheets, customer records, sensor readings—is everywhere in industry. But deep learning models often struggle with it because they treat each number as a single scalar. This paper says, what if we expand each number into a whole vector of spline values first?

Meng: So instead of feeding the model the raw number "forty-two" you're feeding it a richer representation that encodes where "forty-two" falls relative to those knots. That's the encoding part. But my first question as an engineer is always, does this actually help, or is it just adding complexity?

Tom: That's exactly what the authors set out to test. They ran over five thousand training runs across twenty-five datasets, three different backbones, and a whole zoo of encoding methods. They weren't just theorizing; they built a massive benchmark to see what actually works.

Jane: And the title hints at the journey. "From Uniform to Learned Knots" means they started with the simplest approach—evenly spaced knots—and then asked if letting the model move those knots during training, through backpropagation, would give better results. That's the "learned" part.

Lu: It's a natural progression. In classical statistics, free-knot splines were always a pain because optimizing knot locations is a nasty nonconvex problem. But with modern deep learning and differentiable parameterizations, you can make it work. The authors used a softmax-cumsum trick to keep knots ordered while still allowing gradients to flow.

Meng: I like that they kept the downstream models fixed—MLP, ResNet, FT-Transformer—so the only variable was the encoding. That's clean experimental design. It means any performance difference you see is genuinely from the preprocessing, not from some architectural tweak.

Tom: And the results? They're not a simple "one method wins everything." That's what makes this paper interesting. The best encoding depends on whether you're doing regression or classification, which backbone you're using, and even how many basis functions you allow. It's a nuanced picture.

Jane: Which is a great hook for what's coming next. We need to dig into what they actually found, because the summary is where the surprises start.

Summary: Jane: So we've set the stage with the title. Now let's talk about what this paper actually found. Tom, the headline result really surprised me.

Tom: Mine too, Jane. For classification, the humble Piecewise Linear Encoding—PLE—was the most robust choice across the board. It topped the critical difference diagrams at every output size we looked at. The spline methods were competitive, but they didn't consistently beat this simpler baseline.

Jane: But for regression, it was a completely different story. There was no single winner. The best method depended heavily on the output size—how many basis functions you allowed per feature. At small sizes, B-spline variants like BS-LGBM and BS-Q led the pack. At larger sizes, I-splines and the learnable-knot variants started to shine.

Lu: That's a fascinating asymmetry. It suggests that classification tasks, with their threshold-like decision boundaries, are naturally suited to PLE's piecewise-linear structure. Regression targets, on the other hand, are often smoother, and splines—especially integrated splines—can capture that smoothness more efficiently.

Meng: And the backbone mattered too, right? I remember the FT-Transformer was a different beast entirely. For that model, standard scaling was often just as good as any fancy encoding, and sometimes better. That's a practical insight: if you're using a powerful transformer, you might not need to bother with spline preprocessing at all.

Tom: Exactly, Meng. The gains from expressive encodings were much clearer for MLP and ResNet. For FT-Transformer, the benefits were smaller and less consistent. The paper even showed that increasing the output size could hurt FT-Transformer performance in regression, which is a cautionary tale.

Jane: And they didn't just look at performance. They did an efficiency case study on the SGEMM dataset. Learnable knots add parameters, but the real cost was in training time. B-spline learnable knots were relatively cheap, but M-spline and I-spline versions got dramatically slower as the output size grew.

Lu: That makes sense from the math. B-splines have local support—each basis function only depends on a few nearby knots. But I-splines are integrals, so they have cumulative dependence. A small change in one knot can affect many basis functions downstream, which makes the backward pass much more expensive.

Meng: So the practical takeaway for me is: if you want learnable knots, stick with B-splines unless you have a really good reason not to. The performance gains from I-splines might not justify the computational overhead.

Tom: And that's the kind of practical guidance that makes this paper valuable. It's not just "splines are great." It's a map of when and where each encoding strategy pays off. Which brings us to the improvements they suggest—what does this mean for the field going forward?

Improvements: Tom: We've covered the results, but what really excites me is where this paper points the field. Jane, what do you think the biggest improvement is?

Jane: For me, it's the learnable-knot parameterization itself. The authors showed you can optimize knot locations end-to-end in a stable way, using that softmax-cumsum trick to keep them ordered. That's a technical contribution that opens the door for much more adaptive preprocessing.

Lu: And it's a significant step beyond the fixed-knot approaches that dominated before. In classical spline regression, free-knot methods were notoriously fragile. The authors here introduced a collision-avoidance regularization term that penalizes knots getting too close together, which keeps the basis well-conditioned. That's a clever solution to a long-standing problem.

Meng: But let's be practical. The paper also showed that learnable knots can substantially increase training cost, especially for M-splines and I-splines. So the improvement isn't just about making it work—it's about making it work efficiently. The B-spline variant is the sweet spot.

Tom: Right, and the paper also suggests a bigger picture improvement: moving away from a one-size-fits-all approach to numerical encoding. They found that the best method depends on the task, the backbone, and the output size. So the future might be about adaptive selection—choosing the right encoding for each feature, or even learning that choice.

Jane: That's a great point. The authors explicitly mention feature-specific encoding as a future direction. Instead of applying the same spline family and output size to every numerical feature, you could let the model decide which features need more resolution and which ones are fine with a simple scaling.

Lu: And there's room for other basis families too. They mention thin-plate splines and radial basis functions as unexplored territory. Those have different properties—more global support, for example—which might be better suited for certain types of tabular data.

Meng: I'd also love to see this integrated with automated machine learning pipelines. Imagine an AutoML system that not only searches over architectures but also over numerical encodings, using the results from this paper to narrow the search space intelligently.

Tom: That's a compelling vision. The paper gives us a solid foundation—a benchmark, a methodology, and a set of findings—that future work can build on. It's not the end of the story; it's more like the first detailed map of a territory we're just starting to explore.

Jane: And that's exactly what good research should do. It answers some questions but opens up many more. The improvements here are both technical and conceptual, and they're going to shape how people think about tabular deep learning for years to come.

Conclusion: Jane: Well, we've had a great discussion about "From Uniform to Learned Knots: A Study of Spline-Based Numerical Encodings for Tabular Deep Learning." Tom, how would you wrap this up for our listeners?

Tom: I'd say the big message is that numerical encoding is not a trivial preprocessing detail. It's a modeling choice that can have a real impact on performance. The paper shows that spline-based encodings, especially with learnable knots, are a powerful tool—but they're not magic. You have to match the encoding to the task and the backbone.

Jane: And the practical guidance is clear. For classification, PLE is your safest bet. For regression, you have more options, and you should experiment with output size and knot placement. And if you're using a strong transformer like FT-Transformer, don't assume you need fancy encodings at all.

Lu: The methodological contribution is also important. The differentiable knot parameterization with the spacing penalty is a robust way to do free-knot spline optimization in deep learning. That's going to be useful beyond just tabular data—it could apply to any model that uses spline-based components.

Meng: And from an engineering standpoint, the efficiency analysis is invaluable. Knowing that B-spline learnable knots are relatively cheap while I-splines can be expensive helps us make informed trade-offs between accuracy and compute.

Tom: So we're saying goodbye to this paper with a clear picture: it's a comprehensive, well-executed study that gives us both new tools and new understanding. It's the kind of work that will be cited for years as the reference point for numerical encoding in tabular deep learning.

Jane: Absolutely. And it leaves us excited for what's next. Feature-specific encodings, new basis families, integration with AutoML—there's a lot of fertile ground here. We'll be watching for the follow-ups.

Tom: Thanks for joining us, everyone. We'll see you next time with another paper from the arXiv.

Jane: Take care, and keep learning!

Manish Kumar, Anton Frederik Thielmann, Christoph Weisser, Benjamin Säfken

BASF · Clausthal University of Technology · Amazon Music · Bielefeld School of Business, Hochschule Bielefeld

cs.LG

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: 20, 9 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 80/100

The gist: The paper "From Uniform to Learned Knots: A Study of Spline-Based Numerical Encodings for Tabular Deep Learning" by Manish Kumar, Anton Frederik Thielmann, Christoph Weisser, and Benjamin Säfken

Key concepts

Tabular Deep Learning
This field involves using deep learning models to analyze structured data found in spreadsheets or databases. Instead of treating each raw number as a single value, this approach transforms the number into a richer vector representation that captures its position relative to specific bend points or 'knots.'
Piecewise Linear Encoding (PLE)
A simple, robust encoding method where data is represented by piecewise linear segments. The paper found that PLE performs very well across many datasets for classification tasks because its structure aligns naturally with threshold-like decision boundaries.
Learned Knots
This refers to the concept of allowing a deep learning model to optimize or 'learn' the optimal locations of the bend points (knots) during training. This is more adaptive than using fixed, evenly spaced knots, but it increases computational complexity and training time.

Terminology

Summary

The paper From Uniform to Learned Knots: A Study of Spline-Based Numerical Encodings for Tabular Deep Learning by Manish Kumar, Anton Frederik Thielmann, Christoph Weisser, and Benjamin Säfken investigates the role of numerical preprocessing in tabular deep learning, with a particular focus on spline-based numerical encodings.

The authors state that Most tabular datasets contain numerical columns whose effects are often non-uniform. A feature may matter only over specific value ranges, exhibit threshold behavior, or relate to the target through localized changes. However, "a common deep learning pipeline represents each numerical feature as a single scaled scalar, for example, through normalization or min-max scaling, and relies on the backbone to learn nonlinear structure from these inputs. This induces a strong bias toward global smooth transformations and can be mismatched with tabular problems in which predictive structure is tied to specific value ranges."

The authors note that "Prior work shows that the representation of numerical features can substantially affect tabular deep learning performance. In particular, explicit encodings such as piecewise-linear encoding (PLE), and periodic mappings can improve results across several backbones. They also observe that Surveys also note that numerical encodings remain less systematically explored than architectural modifications, despite their practical importance."

The paper studies three spline families for encoding numerical features: B-splines (de Boor, 1972), M-splines (Ramsay, 1988), and integrated splines (I-splines) (Meyer, 2008). Throughout the study, cubic splines with degree p = 3 are used. For each numerical feature xj, the encoding is defined as φj(xj; τj) = (bj,1(xj; τj),..., bj,mj(xj; τj)), where the number of basis functions is determined by mj = Kj + p + 1, with Kj being the number of internal knots.

Four knot-placement strategies are investigated:

  1. Uniform knot placement: Uniform internal knots are equally spaced over the observed range of xj.

  2. Quantile-based knot placement: Quantile internal knots place more knots in regions where samples are concentrated using the empirical quantile function.

  3. Target-aware knot placement: Two variants based on CART and LightGBM split points. The authors state: For each numerical feature xj after min-max scaling to [0, 1], we construct a univariate target-aware set of internal knots by fitting a predictive tree on the training fold only. The procedure involves fitting a depth- and sample-constrained univariate CART tree or a univariate gradient-boosted tree ensemble, collecting split thresholds, applying a minimum spacing constraint, and pruning or supplementing knots to match a target complexity.

  4. Learnable-knot placement: In the learnable-knot variant, also referred to as the gradient-based knot, we treat the internal knots κj as learnable parameters and optimize them jointly with the downstream backbone by backpropagation. The authors use a differentiable parameterization based on ordered spacings, implemented through a softmax followed by cumulative summation, which preserves knot ordering while remaining fully differentiable. A collision avoidance regularization is included: To discourage small interval widths, we penalize the induced spacings using a reciprocal barrier.

The evaluation covers 25 tabular datasets covering regression and classification tasks, collected from the UCI Machine Learning Repository and OpenML. The authors use 5-fold cross-validation for all experiments. In total, this yields 25×5×3×14 = 5250 training runs across 25 datasets, 5 folds, 3 backbones, and 14 numerical encoding methods.

Three backbones are evaluated: MLP, ResNet, and FT-Transformer. The training protocol uses AdamW with learning rate 10−4, weight decay 10−5, batch size 512, and at most 200 epochs with early stopping and a ReduceLROnPlateau scheduler. The per-feature output size is fixed to m ∈ 7, 15, 30 for all features, for both spline encodings and PLE.

The authors find that "The regression results are clearly output-size dependent. At m = 7, the CD diagram in Figure 1 is led by B-spline variants with fixed or target-aware knot placement, with BS-LGBM, BS-Q, and BS-CART occupying the top ranks. At m = 15 and m = 30, the ranking shifts toward I-spline and learnable-knot variants. In particular, IS-Q and IS-LGBM remain among the strongest methods at both larger output sizes, and IS-Grad-U becomes competitive at m = 30."

The heatmaps show that For MLP and ResNet, increasing the output size often improves the average NRMSE of spline-based methods, especially for target-aware and learnable-knot variants. However, "FT-Transformer behaves differently. Several B-spline variants worsen as m increases, for example BS-Q changes from 0.2451 to 0.2666 to 0.3034 across m = 7, 15, 30. Std, with NRMSE 0.2465 remains competitive and in fact outperforms PLE at all three output sizes."

The authors state: "The classification results are more stable than the regression results. In all three CD diagrams in Figure 2, PLE is the top-ranked method, and its average rank improves from 5.0 at m = 7 to 4.0 at m = 15 and 3.3 at m = 30. They further note: Across all three backbones, PLE achieves the highest average AUC at every output size, with small but consistent gains as m increases."

The authors conclude: "The results show that the effect of numerical preprocessing depends on both the task and the backbone. In regression, the strongest methods vary with the output size, with B-spline variants tending to perform best at smaller sizes and I-spline or learnable-knot variants becoming more competitive as the output size increases. In classification, PLE is the most robust choice across backbones and output sizes, while spline-based encodings remain competitive but do not consistently surpass it."

They also note: The heatmaps also indicate that expressive preprocessing is more beneficial for MLP and ResNet than for FT-Transformer, which often shows smaller or less consistent gains, especially in regression.

The authors provide a small controlled illustration of how the two encodings behave when the representation size is held fixed using simple synthetic problems. They find: "In the regression example, the B-spline basis gives a smoother fit and stays closer to the target curve, while the PLE fit shows more visible piecewise-linear changes at the bin boundaries. In the classification example, PLE follows the flat high-probability region and the sharper boundary changes more closely, whereas the B-spline fit changes more gradually across these regions."

The efficiency analysis on the SGEMM dataset reveals: "Among the learnable-knot methods, BS-Grad-U is consistently the cheapest. Its runtime stays relatively stable for MLP and ResNet and increases only moderately for FT-Transformer. By contrast, MS-Grad-U and especially IS-Grad-U become much slower as m increases. The authors explain: B-splines have the most local computation... M-splines add a knot-dependent normalization factor... I-splines inherit this normalization and additionally introduce cumulative dependence through the integral structure."

The ablation study on a synthetic regression task shows: For most methods, test NRMSE improves as m increases from 5 to roughly 15–35, after which the curves mostly plateau. The authors find: Among all configurations, B-spline variants are the strongest overall. The best result is obtained by BS-CART at m = 30 with NRMSE 0.0456 ± 0.0014. They also note: M-spline variants are generally weaker and often deteriorate at larger output sizes, with visibly wider uncertainty bands.

The authors list the following main takeaways:

  • The CD diagrams show statistically significant differences among preprocessing methods across all output sizes for both regression and classification.

  • In regression, the aggregate ranking changes with output size.

  • In classification, PLE is the strongest overall baseline.

  • Larger and more expressive preprocessing tends to benefit MLP and ResNet more than FT-Transformer.

  • Std and MinMax are generally weaker than explicit numerical encodings, especially for MLP and ResNet.

The authors conclude: "In this work, we showed that numerical encoding is an important modeling choice in tabular deep learning rather than a minor preprocessing detail. Our results demonstrate that basis expansion methods, and spline-based encodings in particular, provide a strong alternative to standard scaling approaches and can lead to clear performance gains. We further showed that spline knots can be optimized end to end in a stable manner under the proposed parameterization, making learnable-knot spline encodings a practical preprocessing approach. At the same time, their usefulness depends on the task, backbone, output size, knot-placement strategy, and computational budget, so no single method is uniformly best across all settings."

The authors acknowledge: Our study covers only part of the design space of numerical preprocessing for tabular deep learning. Future work could examine broader adaptive encoding schemes, alternative learnable-knot parameterizations, and additional basis-function families such as thin-plate splines and radial basis functions. They also suggest: A natural extension would be to allow feature-specific choices, with different encodings or encoding sizes assigned to different features.

Improvements for AI systems

Based on the paper, here are specific improvements that can be made to AI systems, along with what the improved system can do:

Improvement: Build an AI system that automatically selects the optimal numerical encoding strategy based on task type, backbone architecture, and desired output size. The paper shows no single encoding dominates across all settings.

What the improved system can do:

  • For classification tasks: Automatically default to PLE (Piecewise Linear Encoding) with output size 30, as it consistently achieves the highest AUC across MLP, ResNet, and FT-Transformer backbones.

  • For regression tasks: Dynamically choose between B-spline variants (at small output sizes m=7) and I-spline variants (at larger output sizes m=15-30), based on the backbone being used.

  • For MLP/ResNet backbones: Prefer larger output sizes (m=30) with learnable-knot splines for regression, as these show the largest gains.

  • For FT-Transformer: Retain standard scaling as a competitive baseline, since explicit encodings show smaller or inconsistent gains here.

These improvements enable AI systems to achieve better predictive performance on tabular data, particularly for regression tasks, while maintaining computational efficiency and training stability.

Sources

Related papers