KANMixer: a minimal KAN-centered mixer for long-term time series forecasting

arXiv:2508.01575 · cs.LG · Submitted 2026-04-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "KANMixer: a minimal KAN-centered mixer for long-term time series forecasting".

Jane: The paper was written by Lingyu Jiang, Dengzhe Hou, Yuping Wang, Yao Su, Shuo Xing et al. from Tohoku University and University of Michigan and Texas A&M University and San Diego State University and Worcester Polytechnic Institute.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv radio hour, everyone. Today we're digging into a paper that's been making the rounds in the forecasting community — it's called "KANMixer: a minimal KAN-centered mixer for long-term time series forecasting." Jane, I've got to say, the title alone tells you exactly what the authors are trying to do — they want to keep it simple, they want to center everything around KAN, and they want to see if that actually works.

Jane: And Tom, for our listeners who haven't been following the KAN wave — KAN stands for Kolmogorov-Arnold Networks. They're a relatively new type of neural network that replaces the fixed activation functions you'd find in a standard MLP with learnable basis functions. Think of it like this: a regular neural network has a set of fixed curves it can use to shape its predictions, but a KAN can actually reshape those curves during training. That flexibility is the whole bet here.

Tom: And the bet seems to be paying off, because the paper reports that KANMixer gets the best MSE in sixteen out of twenty-eight benchmark settings, and best MAE in eleven. That's against nine different baselines, including some pretty heavy hitters like iTransformer and PatchTST.

Jane: Right, and what's striking to me is that they're doing this with a deliberately minimal architecture. No fancy decomposition modules, no patching, no attention mechanisms. Just a multi-scale pooling frontend, a KAN-based temporal mixing backbone, and KAN-based prediction heads. The authors are basically saying — look, we stripped away all the extra machinery, and the KAN core is doing the heavy lifting.

Tom: And that's a bold claim, because the field has been through this cycle before. Remember when DLinear showed that a simple linear model could beat complex Transformers? That shook things up. Now these authors are saying KAN might be the next step in that direction — a more expressive nonlinear core that doesn't need all the hand-crafted priors.

Jane: Exactly. And they're not just claiming it works — they're doing careful ablations to show *why* it works. They found that the prediction head is the most critical component, which is a really useful insight for anyone building forecasting models. But we'll get into that in a bit. First, let's talk about what KAN actually brings to the table here.

Tom: I think the most exciting part is that they're not just adding KAN to an existing architecture — they're building the whole model around it. And the results suggest that the adaptive nonlinearity of KAN is genuinely useful for capturing the kind of non-stationary patterns you see in real-world time series. Electricity demand, weather patterns, traffic — these all have local fluctuations that a fixed-form MLP might miss.

Jane: And that's the hook for our next segment — we're going to look at how they actually designed the architecture and what the main results show. Stay with us.

Summary: Tom: So we're back with "KANMixer: a minimal KAN-centered mixer for long-term time series forecasting." Jane, let's get into the actual architecture and the main results, because I think that's where the excitement really starts.

Jane: Absolutely. So the architecture has three main pieces. First, there's a multi-scale pooling frontend — they take the input and downsample it at multiple resolutions, so the model sees the series at different granularities. Then there's the temporal mixing backbone, which is a stack of KAN-based blocks that process these multi-scale representations. And finally, there are scale-specific KAN prediction heads that produce the forecasts.

Tom: And the key design choice here is that everything is KAN. The mixing, the feed-forward transformations, the prediction heads — all KAN. But they also did component-wise ablations where they replaced each KAN module with an MLP one at a time, and that's where the really interesting finding comes in.

Jane: Right — the prediction head is the most critical component. When they replaced the KAN prediction head with an MLP, the performance dropped the most. That's a really practical insight, because it suggests that if you have an existing forecasting architecture, you might not need to redesign the whole thing — just swapping in a KAN prediction head could give you a meaningful boost.

Tom: And that's the kind of finding that could have real impact, because it's actionable. But let's talk about the main results, because they're pretty impressive. On ETTh1, KANMixer improves MSE by an average of four point nine percent across horizons. On ETTm1 and ETTm2, it leads in three out of four horizon configurations each. And it's doing this with only 321 point 73K parameters.

Jane: That parameter efficiency is worth emphasizing. PatchTST, for comparison, uses three point seven five million parameters. So KANMixer is more than ten times smaller, and still competitive or better on many benchmarks. That's a big deal for deployment scenarios where you're constrained by memory or compute.

Tom: But it's not all roses. On Electricity, which has three hundred twenty-one variables, iTransformer wins because it has explicit cross-variate attention. And on Exchange Rate, Time-FFM benefits from cross-dataset pre-training. So KANMixer's channel-independent design has a clear limitation — it doesn't model relationships between variables.

Jane: And that's an honest limitation to report. The authors acknowledge that KANMixer should be viewed as a diagnostic and reference architecture for studying KAN in LTSF, rather than a production system. But even with that caveat, the results are strong enough to warrant attention.

Tom: And the fact that they're getting these results with a minimal design — no decomposition, no patching, no attention — really raises the question of whether the complexity in existing models is actually necessary. That's a provocative implication.

Jane: It is. And it connects to a broader debate in the field about whether the gains from increasingly complex architectures are real or just artifacts of evaluation settings. KANMixer is a counterpoint — a simple model that performs well because of the expressiveness of its core nonlinearity.

Tom: So what's the catch? What are the trade-offs? That's what we're going to dig into next — the ablations, the basis function choices, and the computational cost.

Improvements: Tom: Welcome back. We're still on "KANMixer: a minimal KAN-centered mixer for long-term time series forecasting." Jane, we've covered the architecture and the main results. Now let's talk about what the paper actually teaches us about how to use KAN effectively — because the ablations are where the real insights are.

Jane: Definitely. And the first big finding is about depth. They compared two three and four-layer KAN backbones, and the three-layer version performed best. The four-layer version was actually less stable — it occasionally diverged during training, though in fewer than five percent of runs. So there's a sweet spot, and it's not deeper.

Tom: That's interesting because in a lot of deep learning, more layers is better. But KAN seems to be different — the optimization gets harder as you stack more KAN layers. The authors suggest this is consistent with prior reports that KAN can be sensitive to initialization and gradient propagation in deeper stacks.

Jane: Right. And then there's the basis function ablation, which I think is one of the most practically useful findings in the paper. They compared B-spline, Chebyshev, Fourier, and Wavelet bases. And B-spline was the only one that consistently outperformed MLP across all prediction lengths. Chebyshev was less consistent, and Fourier and Wavelet both underperformed MLP.

Tom: That's a huge finding, because a lot of the KAN literature has focused on alternative bases for computational efficiency. But this paper shows that the choice of basis function is not a minor implementation detail — it's a primary design variable. The local support of B-splines seems to be what makes them work well for time series.

Jane: And that connects to the theoretical work showing that B-spline KAN is related to radial basis function networks. The local adaptivity is what lets the model capture non-stationary patterns without globally altering the representation. Fourier and Wavelet bases impose more global structure, and that hurts.

Tom: But here's the counterintuitive finding that really got me — decomposition. In a lot of MLP-based models like DLinear and TimeMixer, decomposing the series into trend and seasonal components helps. But in this paper, adding DFT or moving-average decomposition to KANMixer *hurt* performance. It degraded KAN by about zero point zero two five to zero point zero three three MSE, while it actually helped MLP slightly.

Jane: That's a really important result, because it suggests that structural priors developed for MLP-based models don't transfer to KAN-based models. The authors' interpretation is that KAN is already flexible enough to model trend and seasonal structure directly from raw inputs, so imposing decomposition just restricts what it can learn.

Tom: And that's the kind of finding that could change how people design forecasting models. If you're using KAN, you might want to drop the decomposition and let the KAN figure it out. But if you're using MLP, decomposition still helps. It's a genuine interaction between architecture and prior.

Jane: And then there's the multi-scale design. Removing it hurt KAN performance by about zero point zero two zero MSE. So KAN benefits from multi-resolution input enrichment, as long as that enrichment doesn't impose rigid structural constraints. That's a useful distinction — enrichment is good, decomposition is bad.

Tom: So the practical guidance is: use B-spline KAN, use a moderate depth of three blocks, put KAN in the prediction head, and skip the decomposition. That's a pretty actionable set of recommendations.

Jane: It is. But there's also the computational cost question. And that's where the paper is honest about a real bottleneck. Let's talk about that.

Conclusion: Tom: And we're wrapping up our discussion of "KANMixer: a minimal KAN-centered mixer for long-term time series forecasting." Jane, we've covered the architecture, the results, and the ablations. Let's pull it all together.

Jane: So the big picture is this — KANMixer shows that a minimal KAN-centered architecture can be competitive with much more complex models on long-term time series forecasting. It achieves the best MSE in sixteen of twenty-eight settings, and it does so with a fraction of the parameters of models like PatchTST.

Tom: And the ablations give us concrete guidance. B-spline bases are the way to go. Three blocks is the sweet spot for depth. The prediction head is where KAN matters most. And decomposition — which helps MLP — actually hurts KAN. That last finding is genuinely surprising and could reshape how people think about structural priors in forecasting.

Jane: But there are real limitations. The computational cost is a bottleneck — the B-spline variant takes about three times longer to train per epoch than the MLP baseline, largely because spline evaluation lacks optimized kernels. And the channel-independent design means it doesn't capture cross-variable relationships, which is why it loses on Electricity.

Tom: And the authors are upfront that this should be viewed as a diagnostic architecture for studying KAN, not a production system. But the insights are transferable — especially the finding about the prediction head. If you're building a forecasting model, swapping in a KAN prediction head might be a cheap win.

Jane: I think the broader implication is that the field has been layering on complexity — decomposition, patching, attention — and this paper suggests that a more expressive core nonlinearity might make some of that complexity unnecessary. That's a provocative idea, and it's one that deserves more investigation.

Tom: And the future work is clear — extending KANMixer with lightweight cross-variate modeling, developing better spline kernels, and testing whether the prediction-head finding transfers to Transformer-based models. There's a lot to build on here.

Jane: Absolutely. So we'll say goodbye to KANMixer and get ready for the next paper. Thanks for listening, everyone.

Tom: And remember — the paper is "KANMixer: a minimal KAN-centered mixer for long-term time series forecasting." If you're working on forecasting, this one's worth a read. See you next time.

Lingyu Jiang, Dengzhe Hou, Yuping Wang, Yao Su, Shuo Xing, Wenjing Chen, Xin Zhang, Zhengzhong Tu, Ziming Zhang, Fangzhou Lin, Michael Zielewski, Kazunori D Yamada

Tohoku University · University of Michigan · Texas A&M University · San Diego State University · Worcester Polytechnic Institute

cs.LG

Submitted: 2026-04-22

Updated: 2026-08-18

Comments: 11 pages, 3 figures, 5 tables

Code: https://github.com/thuml/Time-Series-Library

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 60/100

The gist: a multi-scale pooling frontend, a KAN-based temporal mixing backbone, and KAN-based prediction heads.

Key concepts

Kolmogorov-Arnold Networks (KAN)
A type of neural network that replaces fixed activation functions found in standard MLPs with learnable basis functions. This allows the KAN to reshape curves during training, providing increased flexibility for modeling complex patterns.
Time Series Forecasting
The process of predicting future values based on historical data points. The episode discusses applying advanced models like KANMixer to predict variables such as electricity demand or weather patterns over long periods.
B-spline Basis
One of the basis functions tested in the paper. B-splines were found to be the most effective choice for KAN in time series, suggesting that their local support helps capture non-stationary patterns effectively.

Terminology

Summary

Summary

The paper introduces KANMixer, a minimal Kolmogorov-Arnold Network (KAN)-centered architecture for long-term time series forecasting (LTSF). The authors investigate whether KANs, which feature adaptive basis functions capable of granular modulation of nonlinearities, can improve LTSF performance and under which design choices they are most effective. The proposed architecture consists of three modules: a multi-scale pooling frontend, a KAN-based temporal mixing backbone, and KAN-based prediction heads. The design deliberately avoids decomposition and other heavy engineering modules, enabling controlled attribution of performance gains to KAN components.

Main results. Across 28 benchmark–horizon settings against nine baselines, KANMixer achieves the best MSE in 16 settings and the best MAE in 11. In controlled comparisons (KANMixer, TimeKAN, and TimeMixer, all reproduced in the same environment with identical splits, RevIN normalization, and evaluation code, averaged over five runs), KANMixer achieves lower MSE than both TimeKAN and TimeMixer in 14 of 28 settings, with mean differences of 0.010–0.020 MSE on ETTh1 and 0.005–0.015 on ETTm1/ETTm2, consistently exceeding within-model standard deviations (≤0.005 MSE). The advantage is most pronounced on ETTh1, where KANMixer improves MSE by an average of 4.9% across horizons. On ETTm1 and ETTm2, KANMixer leads in 3 of 4 horizon configurations each. On Electricity (321 variables), iTransformer leads due to explicit cross-variate attention; on Exchange Rate, Time-FFM benefits from cross-dataset pre-training. KANMixer outperforms Time-FFM on 5 of 7 datasets using only 321.73K parameters.

Ablation studies. Five ablation studies were conducted on three representative datasets (ETTh1, ETTm1, Weather):

  1. Depth ablation (KAN vs. MLP): Replacing KAN layers with MLP layers shows KAN consistently outperforms MLP across all three datasets. The 3-layer KAN variant gives the best overall performance. Deeper KAN stacks (4 layers) exhibit mild instability, with occasional divergence in fewer than 5% of runs, not observed for KAN-3L.

  2. Component-wise ablation: Replacing each KAN module with its MLP counterpart shows the KAN prediction head is the most critical component — replacing it leads to the largest degradation among all module-wise substitutions. The gains from KAN-based mixing and feed-forward transformation are comparatively smaller.

  3. Basis function ablation: Comparing four basis function variants (B-spline, Chebyshev, Fourier, Wavelet) with an MLP baseline, B-spline is the only basis that consistently outperforms MLP across all prediction lengths. Chebyshev performs less consistently and degrades at longer horizons on ETTh1. Fourier and Wavelet both underperform MLP across most settings, with Wavelet showing pronounced instability at longer horizons.

  4. Structural prior and multi-scale ablation: Introducing DFT and moving-average (MA) decomposition, or removing the multi-scale module (NoMS), consistently hurts the KAN variant while producing only marginal gains or mixed effects for the MLP counterpart. Decomposition-based priors lead to clear error increases for KAN (+0.025 to +0.033 MSE) while yielding small improvements for MLP (−0.003 to −0.011 MSE) under the same protocol. Removing the multi-scale module degrades KAN performance by about 0.020 MSE.

Computational efficiency. KANMixer remains highly compact: the MLP baseline uses 92.9K parameters, while Chebyshev, Wavelet, B-spline, and Fourier variants use 160.9K, 120.7K, 321.7K, and 371.1K parameters, respectively — all substantially smaller than PatchTST (3.75M parameters). Arithmetic cost varies across basis functions: Chebyshev KANMixer requires 22.93M MACs versus 21.44M for MLP, while Wavelet, B-spline, and Fourier increase cost to 39.12M, 90.57M, and 126.04M MACs, respectively, still far lighter than PatchTST at 5.89G MACs. GPU memory averaged across prediction lengths: MLP baseline uses 1.6 GiB, B-spline and Fourier KANMixer require 7.5 GiB and 8.7 GiB, far below iTransformer's 38.7 GiB. At P=720, B-spline and Fourier require 13.7 GiB and 15.8 GiB, compared with 2.3 GiB for MLP and 40.9 GiB for iTransformer. Training time shows an implementation bottleneck: Chebyshev KANMixer requires 29.4 s/epoch versus 16.6 s/epoch for MLP despite nearly the same MAC count; B-spline is slowest at 49.8 s/epoch on average and 51.5 s/epoch at P=720, while PatchTST trains in 7.1 s/epoch on average.

Key findings and interpretations. The paper reports that KAN is most effective when the basis function is chosen appropriately, when the architecture allows KAN to learn from enriched but unconstrained inputs, and when adaptive nonlinearity is preserved at the prediction head rather than replaced by a fixed-form MLP. The authors note that decomposition, which consistently helps MLP, degrades KAN performance, suggesting that structural priors developed for MLP-based forecasting may not transfer directly to KAN-based architectures. The paper proposes that B-spline functions' local support and adaptive knot placement may help capture non-stationary temporal patterns, consistent with the theoretical characterization of B-splines as locally supported piecewise polynomial functions. The authors also suggest that the prediction-head finding may help explain TimeKAN's observation that KAN and MLP behave similarly within its frequency decomposition pipeline, since placing KAN mainly in intermediate layers while keeping an MLP prediction head could attenuate overall gain.

Limitations. The ablation findings are established on three datasets under a single experimental protocol and should not be assumed to generalize directly to strongly cross-correlated high-dimensional multivariate datasets, settings with periodic structure beyond the look-back window, more severe distribution shifts, or substantially longer prediction horizons. KANMixer adopts a channel-independent design and does not explicitly model cross-variable dependencies, placing it at a disadvantage on datasets where cross-variate interactions are important. The explanations for why decomposition degrades KAN and why the prediction head contributes most are empirically motivated interpretations rather than formally established mechanisms. B-spline KAN incurs substantial training-time overhead relative to MLP, approximately threefold per epoch, largely because spline evaluation lacks optimized CUDA kernels. The main comparison table combines three reproduced models with seven literature-reported results, so these two types of comparison should be interpreted differently.

Methodology. KANMixer is a channel-independent multi-scale forecasting architecture. Given an input sequence of shape (B, L, C), it constructs a list of downsampled sequences at multiple temporal resolutions (default k=3 scales with downsampling windows of 2, giving lengths approximately L, L/2, L/4). Each scale is normalized independently using reversible instance normalization (RevIN), embedded into a latent representation of shape (BC, Li, dmodel) via a position-free token embedding layer, processed jointly by a stack of N=3 temporal mixing blocks, and each scale produces its own forecast through scale-specific KAN prediction heads, summed across scales before inverse normalization. Each temporal mixing block performs fine-to-coarse temporal mixing across scales (using KAN down-mixing modules) and within-scale feature refinement (using residual KAN feed-forward transformations). KAN layers use B-spline basis functions with grid size G=5 and spline order p=3 by default, with each edge function parameterized as the sum of a base branch (SiLU activation) and a spline branch. The latent dimension is dmodel=16, the feed-forward hidden width is dff=32, and total parameter count at default settings is 321.73K. Training uses the Adam optimizer with learning rate 0.001 and cosine annealing decay, batch size 32, MSE training loss, gradient clipping at max norm 1.0, early stopping with patience of 3 epochs based on validation MSE, and training capped at 10 total epochs. A divergence criterion is applied: if training loss at any epoch exceeds 10× the epoch-1 loss, the run is discarded and restarted with a new random seed. Evaluation follows the established LTSF protocol with look-back window L=96 and prediction lengths P ∈ 96, 192, 336, 720, reporting MSE and MAE. ETT datasets use 6:2:2 train/validation/test splits; Exchange Rate, Weather, and Electricity use 7:1:2 splits.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: In existing LTSF architectures (Transformers, MLPs), replace only the final prediction head with a B-spline KAN layer (grid size 5, order 3), keeping the rest of the architecture unchanged.

What the improved system can do: Achieve the largest accuracy gain per modification—the paper shows this single change yields the most significant MSE reduction compared to replacing any other module. This is a low-cost, high-impact upgrade for any forecasting model.

Improvement: When integrating KAN into time-series models, select B-spline basis functions (local support, adaptive knot placement) over Fourier, Wavelet, or Chebyshev alternatives.

What the improved system can do: Consistently outperform MLP baselines across all prediction horizons (96–720 steps). Fourier and Wavelet variants underperform even plain MLP, and Chebyshev degrades at longer horizons. B-spline’s local support captures non-stationary patterns and abrupt changes that global bases miss.

Improvement: In KAN-based forecasting architectures, eliminate DFT or moving-average decomposition modules that are standard in MLP-based models (e.g., DLinear, Autoformer, TimeMixer).

What the improved system can do: Avoid the counterintuitive performance degradation observed (+0.025 to +0.033 MSE) when decomposition is added to KAN. The system learns frequency and trend structure directly from raw inputs using adaptive basis functions, without needing pre-filtered components that restrict its flexibility.

Improvement: Configure KAN-based temporal mixing backbones with exactly 3 blocks, not 2 or 4.

What the improved system can do: Achieve the best MSE across ETTh1, ETTm1, and Weather datasets. Deeper stacks (4 layers) introduce training instability (divergence in <5% of runs) with no accuracy benefit, while shallower stacks (2 layers) underperform. This depth is robust and stable.

Improvement: Enrich inputs with multi-resolution average pooling (3 scales) before KAN processing, without imposing rigid decomposition.

What the improved system can do: Improve MSE by 0.020 compared to single-scale inputs. The system benefits from multi-resolution information while preserving original signal structure, enabling KAN to exploit locally adaptive features across different temporal granularities.

Improvement: Process each time-series variable independently with shared model weights, rather than jointly modeling cross-variate dependencies.

What the improved system can do: Achieve state-of-the-art results on univariate and low-dimensional multivariate datasets (ETTh1, ETTm1, ETTm2, Weather) with only 321.73K parameters—substantially smaller than PatchTST (3.75M) and more memory-efficient than iTransformer (38.7 GiB vs. 7.5 GiB peak GPU memory). This is ideal for resource-constrained deployment.

Improvement: Set the latent embedding dimension to 16 for KAN-based models, rather than larger values (32, 64, 128) that benefit MLP variants.

What the improved system can do: Achieve peak performance with minimal parameters. KAN’s adaptive basis functions compensate for lower dimensionality, while MLP requires wider layers to match performance. This reduces memory and computation without sacrificing accuracy.

Improvement: Apply gradient clipping (max norm 1.0) and a restart protocol: if training loss exceeds 10× the epoch-1 loss, discard the run and restart with a new seed.

What the improved system can do: Ensure stable convergence in 95%+ of runs, even for deeper KAN stacks. This makes KAN-based systems reliable for production use, where training failures are costly.

Improvement: Develop or integrate CUDA-optimized kernels for B-spline basis function evaluation, which is currently the primary training bottleneck (49.8 s/epoch vs. 16.6 s/epoch for MLP).

What the improved system can do: Reduce training time by up to 3× without changing model accuracy. This addresses the main practical limitation of KAN-based systems, making them viable for large-scale, time-sensitive forecasting tasks.

Improvement: Sum scale-specific forecasts directly (no learned weighting) after KAN prediction heads.

What the improved system can do: Maintain simplicity and avoid overfitting from extra fusion parameters. The system achieves competitive results (best MSE in 16 of 28 benchmark–horizon settings) while keeping the architecture minimal and interpretable, facilitating controlled attribution of performance gains.

Sources

Related papers