SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions".
Jane: The paper was written by Hoang-Thang Ta from University of Information Technology and Vietnam National University Ho Chi Minh City.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We're starting things off with a fascinating new paper called SechKAN: Kolmogorov–Arnold Networks with Hyperbolic Secant Functions.
Jane: It has such a heavy mathematical name, but it’s actually an quite elegant idea once you peel back the layers.
Tom: You're talking about how the author, Hoang-Thang Ta from Vietnam, is trying to fix some of the big headaches we have with current neural networks?
Jane: Exactly, because for a long time we've relied on these standard multilayer perceptrons that can sometimes struggle with very smooth or complex patterns.
Tom: And this new SechKAN approach uses something called Kolmogorov-Arnold Networks to try and change that fundamental structure.
Jane: It’s like moving from using rigid, straight lines to using more flexible, natural curves to describe the world.
Lu: I find that part so beautiful because it suggests we're moving toward a math that actually mimics how things move in nature.
Tom: Do you think that's why people are getting so excited about these KAN architectures lately, Lu?
Lu: Definitely, since instead of just stacking layers of weights, we're giving the network actual functions it can learn and shape.
Meng: That sounds great in a lab, but I have to wonder if this is actually going to be a nightmare for us to deploy on standard GPUs.
Tom: You're thinking about the implementation side of things, Meng?
Meng: Yeah, because many of these new KAN variants use really complex math like B-splines that are hard to optimize and slow down our training pipelines.
Jane: That's a valid concern, but I think this paper is trying to address exactly that kind of complexity.
Lalam: If they succeed in making these functions easier to compute, we could see intelligence becoming much more lightweight and widely distributed.
Tom: That's a big vision to hold onto as we move into the actual mechanics of how SechKAN works.
Summary: Tom: We're continuing our discussion on SechKAN: Kolmogorov–Arnold Networks with Hyperbolic Secant Functions, specifically looking at the "sech" part of the name.
Jane: The hyperbolic secant function is really the star of the show here because it has this lovely bell shape.
Tom: And that shape isn't just for looks, is it?
Jane: Not at all, because that bell shape is very smooth and stays localized, meaning the network can focus on specific data points without everything getting messy.
Tom: So it avoids that problem where one part of the network accidentally messes up what another part is learning?
Jane: That's a good way to put it, as it provides much more stable gradients during training.
Tom: But how does he keep the model from becoming absolutely massive in terms of parameters?
Jane: He uses a clever trick called a 1D linear projection to act as a structural bottleneck.
Tom: So instead of having a huge number of parameters for every single connection, he squeezes them down?
Jane: Exactly, which allows SechKAN to stay roughly the same size as a standard MLP while still being much more expressive.
Lu: I love how that creates this sense of fluid interaction between the layers rather than just heavy-handed connections.
Tom: Do you think that bottleneck might actually limit what the network can learn, Lu?
Lu: It's a balance, but by using those smooth sech functions, the information being squeezed is much higher quality to begin with.
Meng: I'm still a bit skeptical about whether that squeeze loses too much detail for high-precision tasks.
Meng: If the projection is too aggressive, we might just end up with a very efficient way to lose data.
Jane: That's why the results are so important, Meng, because they show the model keeps its power despite being smaller.
Lalam: This is essentially teaching machines to prioritize meaning over raw volume.
Lalam: We are seeing a shift toward quality over quantity in model architecture.
Tom: We should definitely look at those performance numbers next to see if that intuition holds up.
Improvements: Tom: We've talked about the math and the structure of SechKAN: Kolmogorov–Arnold Networks with Hyperbolic Secant Functions, so let's talk about how it actually performs.
Jane: The results are quite striking, especially when you look at how well it handles complex physics simulations.
Tom: You mean things like the Navier-Stokes equations for fluid dynamics?
Jane: Yes, and it even performed really well on the Shallow Water problem, showing it can handle real-world scientific data.
Tom: And it wasn't just limited to science; it also did a great job on image classification benchmarks like CIFAR-one hundred.
Jane: It actually outperformed several other KAN variants in those image tasks, which is no small feat.
Tom: But there is a catch, isn't there?
Jane: There is, because while the parameter count is low, the actual computational cost in terms of FLOPs and memory usage is higher than a standard MLP.
Tom: So it's like having a very smart brain that just needs a bit more energy to run?
Jane: That's an excellent analogy, as it's trading speed for that extra layer of mathematical sophistication.
Meng: That’s the reality of engineering, because if the training throughput drops too much, it becomes hard to justify in a production environment.
Tom: Do you think the accuracy gains are enough to make up for that extra time, Meng?
Meng: It depends on whether you're doing general consumer AI or high-stakes scientific modeling where precision is the priority.
Lu: For researchers, this could be a total game-changer because it opens up new ways to model continuous physical systems.
Tom: You see it as a tool for the next generation of scientific discovery, Lu?
Lu: Absolutely, because we're finally moving past rigid approximations and toward something that understands the smoothness of reality.
Lalam: It also suggests that we can build much more specialized tools for different scientific disciplines rather than relying on one-size-fits-all models.
Lalam: This is a move toward more reliable and specialized AI tools for scientists.
Tom: That's a huge thought to leave us on as we wrap this all up.
Conclusion: Tom: We have covered so much ground today, from the bell-shaped functions to those intense physics simulations in SechKAN: Kolmogorov–Arnold Networks with Hyperbolic Secant Functions.
Jane: It’s been a fascinating look at how changing a single mathematical basis can shift the entire paradigm of network design.
Tom: Before we head out, I want to hear one last thought from our guests.
Lu: My mind is racing with the possibilities of using these smooth functions in robotics and motion planning to create lifelike movement.
Meng: And I'll be digging into those FLOP counts to see how we can optimize this for production-ready hardware.
Lalam: I believe this marks a step toward an AI culture that values elegance and efficiency over sheer scale.
Tom: Well, thank you all for joining us on the show today.
Jane: We'll see you next time when we uncover another incredible paper!
Tom: Goodbye, everyone!
University of Information Technology · Vietnam National University Ho Chi Minh City
cs.LG, cs.AI
Submitted: 2026-06-30
Updated: 2026-09-24
Comments: 37 pages
Code: https://github.com/hoangthangta/All-KAN
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 74/100
The gist: SechKAN is a novel Kolmogorov–Arnold Network (KAN) architecture that utilizes hyperbolic secant (sech) functions as its basis.
Key concepts
- Kolmogorov-Arnold Networks
- Instead of stacking layers of weights like standard neural networks, these architectures provide the network with learnable functions. This allows the model to use flexible, natural curves to describe data rather than the rigid, straight lines used in traditional multilayer perceptrons.
- Hyperbolic Secant (sech) Function
- This function features a smooth, bell-shaped curve that stays localized. This shape allows the network to focus on specific data points without causing interference between different parts of the model, which provides more stable gradients during the training process.
- 1D Linear Projection
- This technique acts as a structural bottleneck to prevent the model from having an excessive number of parameters. By squeezing the connections, it allows SechKAN to remain roughly the same size as a standard multilayer perceptron while remaining highly expressive.
Terminology
Summary
SechKAN is a novel Kolmogorov–Arnold Network (KAN) architecture that utilizes hyperbolic secant (sech) functions as its basis. It addresses the critical limitations of traditional KANs, specifically their parameter inefficiency
and high computational overhead, by providing a parameter-efficient nonlinear representation learning
framework that maintains a model size comparable to that of multilayer perceptrons (MLPs).
Architecture and Basis Design
The core innovation lies in adopting the hyperbolic secant basis, which provides a smooth, bell-shaped form, localized responses, and well-behaved gradients.
Unlike the B-splines used in other KANs, the sech function is infinitely differentiable
with an analytic closed form
and exhibits smooth exponential decay.
This mathematical property promotes stable training and efficient grid aggregation.
The SechKAN layer processes data through a structured sequence:
-
An initial normalization step (Norm1) to rescale input to the grid range [g, g].
-
A hyperbolic secant basis expansion (X') using a learnable grid C, scale s, bias delta, and width w.
-
A
grid-wise linear projection
that compresses the G basis responses into a single scalar (G to 1). -
A second normalization (Norm2) followed by an element-wise nonlinear activation.
-
A feature-wise linear projection to produce the final output, with an optional
skip projection branch.
Parameter Reduction Strategy
To mitigate the high parameter complexity inherent in traditional KANs, SechKAN introduces a lightweight 1D linear projection as a structural bottleneck.
In standard KANs, each edge learns an independent expansion; however, SechKAN shares the nonlinear feature representation across all output neurons before applying linear mixing.
This decomposes the mapping into two stages: the learnable projection of basis responses within each input feature and the subsequent linear mixing across input dimensions.
This factorization ensures that when the grid size G is small relative to the layer dimensions, SechKAN's parameter count is only slightly higher than an MLP. Specifically, it introduces only a small number of additional parameters,
making its complexity comparable to an MLP while retaining the nonlinear basis representation.
Experimental Results and Efficiency
The paper demonstrates SechKAN's effectiveness across three primary tasks:
-
Function fitting: It achieves performance comparable to MLPs and KAN variants, excelling at approximating functions with
oscillatory patterns, multiplicative interactions, and higher-dimensional nonlinear coupling.
-
PDE surrogate modeling: On Navier–Stokes and Shallow Water datasets, it outperforms MLPs and achieves
competitive or better performance than representative KAN variants.
-
Image classification: Across MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100, SechKAN
achieves the best performance among the evaluated KAN variants
while remaining competitive with MLPs.
While SechKAN offers superior predictive power in several benchmarks, the authors note a trade-off: it still incurs higher computational cost than MLPs and some KAN variants.
The model's efficiency is task-dependent, and its performance is highly sensitive to grid size, activation functions, and normalization strategies.
Improvements for AI systems
Improvement 1: Integration of SechKAN layers into Physics-Informed Neural Networks (PINNs) and PDE Surrogate Models.
- Improved AI System Capability: The system can model complex fluid dynamics (e.g., Navier-Stokes) and geophysical wave propagation (e.g., Shallow Water equations) with significantly higher numerical fidelity (lower Relative L2 error) and greater training stability compared to B-spline-based KANs or standard MLPs.
Improvement 2: Deployment of SechKAN as a high-precision regression backbone for complex signal processing and scientific computing.
- Improved AI System Capability: The system can approximate highly nonlinear, oscillatory, and frequency-modulated functions (e.g., those featuring multiplicative interactions and anisotropic patterns) with superior Mean Squared Error (MSE) performance, specifically excelling in high-dimensional variable interactions where traditional MLPs fail.
Improvement 3: Replacement of MLP classification heads in Convolutional Neural Networks (CNNs) with SechKAN heads (SechKAN CNN).
- Improved AI System Capability: The system can achieve higher validation accuracy and macro F1-scores on complex image classification tasks (e.g., CIFAR-10 and CIFAR-100) while simultaneously mitigating the overfitting observed in standard CNN architectures.
Improvement 4: Implementation of SechKAN in parameter-constrained environments (Edge AI/Mobile Deployment).
- Improved AI System Capability: The system can leverage the superior nonlinear representation power of Kolmogorov-Arnold Networks while maintaining a model size and parameter count comparable to traditional Multilayer Perceptrons (MLPs), enabling advanced nonlinear modeling on hardware with strict memory and storage limitations.
Improvement 5: Utilization of learnable hyperbolic secant width (theta w) and 1D grid projection for deep architectural optimization.
- Improved AI System Capability: The system can maintain stable gradient flow and prevent activation saturation in deep, highly nonlinear networks, ensuring more reliable convergence and reducing sensitivity to random initialization during the training of complex, multi-layered models.
Abstract
In recent years KolmogorovArnold Networks KANs have attracted increasing attention due to their effectiveness in machine learning and scientific computing offering a new paradigm for neural network design In this paper we present SechKAN a novel KAN based on hyperbolic secant sech functions The hyperbolic secant basis is adopted for its smooth bellshaped form localized responses and wellbehaved gradients We employ a 1D linear projection to reduce the number of parameters allowing SechKAN to maintain a model size comparable to that of multilayer perceptrons MLPs Experimental results show the effectiveness of SechKAN on function fitting PDE surrogate modeling and image classification benchmarks including MNIST FashionMNIST CIFAR10 and CIFAR100 On function fitting SechKAN achieves performance comparable to both MLPs and representative KAN variants On PDE surrogate modeling it outperforms MLPs and achieves competitive or better performance than representative KAN variants On image classification benchmarks SechKAN achieves the best performance among the evaluated KAN variants while remaining competitive with MLPs using a comparable number of parameters However SechKAN still incurs higher computational cost than MLPs and some KAN variants Our source code is publicly available at httpsgithubcomhoangthangtaAllKAN.
Sources
- KASAM: Spline Additive Models for Function Approximation
- A Temporal Kolmogorov-Arnold Transformer for Time Series Forecasting
- Kolmogorov-Arnold Networks for Time Series: Bridging Predictive Power and Interpretability
- Demonstrating the Efficacy of Kolmogorov-Arnold Networks in Vision Tasks
- Chebyshev Polynomial-Based Kolmogorov-Arnold Networks: An Efficient Architecture for Nonlinear Function Approximation
- Kolmogorov-Arnold Networks are Radial Basis Function Networks
- Architectural Scaling Surpass Basis Complexity? Efficient KANs with Single-Parameter Design
- Enhancing Graph Collaborative Filtering with FourierKAN Feature Transformation
- Wav-KAN: Wavelet Kolmogorov-Arnold Networks
- Unveiling the Power of Wavelets: A Wavelet-based Kolmogorov-Arnold Network for Hyperspectral Image Classification
- PRKAN: Parameter-Reduced Kolmogorov-Arnold Networks
- GS-KAN: Parameter-Efficient Kolmogorov-Arnold Networks via Sprecher-Type Shared Basis Functions
- GroupKAN: Efficient Kolmogorov-Arnold Networks via Grouped Spline Modeling
- KAN or MLP: A Fairer Comparison
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks