HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks
Zhao Su, Yuxin Xia, Haoran Li, Jun Shen, Qi Zhu, Qingguo Zhou, Binbin Yong
Lanzhou University · Monash University · University of Wollongong · Nanjing University of Aeronautics and Astronautics
cs.LG, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks Abstract Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights
Terminology
Summary
HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks
Abstract
Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions. However, assigning an independent function to every connection results in substantial parameter redundancy, limiting their scalability and efficiency. To reduce this redundancy, we introduce HYperbolic Dynamic Representation Architecture (HYDRA), a parameter-efficient hyperbolic extension of KAN that combines spline-based functional learning with representations in the Poincaré ball. HYDRA maps vector-valued inputs into a bounded hyperbolic latent space, performs KAN-style updates in tangent space, and employs a low-rank prototype block to share functional transformations across hidden dimensions. The resulting hyperbolic representations provide a structured radial coordinate for interpretation, while radius control improves training stability by preventing boundary saturation. Extensive experiments across eight benchmark datasets demonstrate that HYDRA consistently achieves competitive or superior predictive performance while improving parameter efficiency and representation interpretability.
Introduction
Many supervised learning problems with vector-valued inputs require expressive nonlinear transformations while keeping the learned model compact and inspectable. Multilayer perceptrons (MLPs) provide flexible approximation but hide feature transformations inside dense scalar weights. Kolmogorov-Arnold Networks (KANs) replace scalar weights with learnable univariate functions on edges (Liu et al. 2025), improving local functional expressiveness but making a dense hidden-to-hidden block scale quadratically with hidden width and linearly with the number of spline bases.
The parameter-efficiency issue is closely related to how hidden representations are organized. Standard KANs operate in Euclidean spaces, where achieving stronger separation or richer representations often requires additional hidden dimensions or spline structures, increasing the parameter count. By organizing representations in a hyperbolic latent space, the model can encode variations more compactly without substantially enlarging the KAN architecture. Meanwhile, the KAN computation is performed in the tangent space, preserving the simplicity and low-parameter nature of Euclidean spline operations. Hyperbolic representation learning provides a natural alternative because distances and volumes grow differently from Euclidean spaces (Nickel and Kiela 2017; Ganea, Bécigneul, and Hofmann 2018b). However, naive hyperbolic modeling creates a new problem: near the Poincaré-ball boundary, distances, tangent coordinates, and gradients are amplified, so a model may separate data by drifting outward rather than by learning stable functional structure.
We propose HYperbolic Dynamic Representation Architecture (HYDRA), a radius-constrained low-rank neural network architecture that integrates hyperbolic representation learning with KAN-based function approximation. The key idea is to decouple representation scale from local functional modeling. Specifically, the Poincaré radius serves as a compact coordinate for encoding latent magnitude variations, while KAN splines capture local functional responses in the tangent space. HYDRA first maps normalized inputs into a bounded Poincaré ball, performs KAN-style residual updates in tangent coordinates, compresses these updates through a low-rank prototype bottleneck, and explicitly regulates the latent radius to prevent uncontrolled geometric expansion. Rather than assuming hierarchical structure in the input domain, HYDRA leverages hyperbolic geometry as an efficient and controllable representation space, while preserving the low-parameter and computationally efficient characteristics of KAN computation.
This design leads to the following contributions:
• We propose HYDRA, a hyperbolic functional learning architecture composed of multiple HYDRA blocks, which perform spline-based KAN updates in tangent space while maintaining bounded Poincaré representations.
• HYDRA introduces a low-rank prototype KAN update that reduces the dominant hidden-to-hidden parameter complexity from O(d2K) to O(dr + r2K).
• HYDRA incorporates a radius-control mechanism to constrain hyperbolic representations and mitigate unstable near-boundary effects.
• Experiments on eight datasets demonstrate that HYDRA achieves competitive or superior predictive performance while using fewer parameters than existing approaches.
Related Work
Kolmogorov-Arnold Networks: KANs have renewed interest in neural architectures where nonlinear transformations are represented as explicit functions rather than implicitly encoded through dense scalar weights. In the original formulation, each edge carries a learnable univariate function, typically parameterized by splines, providing a more interpretable alternative to conventional multilayer perceptrons (Liu et al. 2025). This formulation has motivated subsequent studies on KANs for scientific machine learning and broader neural architectures (Liu et al. 2024; Somvanshi et al. 2025). Despite their interpretability advantages, dense KAN layers suffer from rapidly increasing functional parameters because each input-output connection requires an independent function. Recent variants address this limitation by modifying functional bases or introducing parameter-sharing mechanisms. Chebyshev KAN replaces spline functions with polynomial bases (Sidharth et al. 2024), Wavelet KAN introduces wavelet-based representations (Bozorgasl and Chen 2024), while FastKAN, radial basis function KAN, and parameter-reduced KAN variants further improve efficiency through alternative parameterizations (Li 2024; Ta et al. 2025). These methods reduce functional representation costs while maintaining explicit function learning. However, existing KAN variants mainly focus on edge-function parameterization, leaving the role of latent representation geometry largely unexplored. HYDRA extends this line of research by coupling KAN-based functional updates with structured hyperbolic representations.
Hyperbolic Representation Learning: Hyperbolic representation learning studies negatively curved spaces for efficient representation of complex structures. Early studies introduced Poincaré and Lorentz embeddings, demonstrating the effectiveness of hyperbolic spaces for hierarchical representation learning (Nickel and Kiela 2017, 2018). These ideas were later extended to neural computation through hyperbolic entailment regions, hyperbolic neural networks, and Riemannian optimization methods (Ganea, Bécigneul, and Hofmann 2018a,b; Bécigneul and Ganea 2019). Recent research has expanded hyperbolic learning beyond embedding problems toward complete neural architectures, including fully hyperbolic networks, hyperbolic graph models, attention mechanisms, Transformers, supervised representation learning, and residual architectures (Shimizu, Mukuta, and Harada 2021; Chen et al. 2022; Yang et al. 2023, 2024; Nock et al. 2024; Sinha et al. 2024; Li et al. 2024; He, Yang, and Ying 2025). More recently, hyperbolic geometry has also been explored for large-model adaptation and foundation-model learning (Yang et al. 2025; He et al. 2025). However, existing hyperbolic architectures mainly focus on geometric representation learning or neural computation in curved spaces, while the interaction between hyperbolic representations and explicit functional learning remains less explored. HYDRA explores this direction by performing KAN-based functional updates in tangent spaces associated with bounded hyperbolic representations.
Interpretable Neural Networks: Interpretability research aims to understand how neural models transform inputs into predictions. Generalized additive models and their neural extensions improve transparency by explicitly modeling feature-wise effects and interactions (Yang, Zhang, and Sudjianto 2021; Agarwal et al. 2021; Chang, Caruana, and Goldenberg 2022). Meanwhile, SHAP values provide a widely adopted model-agnostic framework for estimating feature contributions (Lundberg and Lee 2017). KANs further introduce intrinsic interpretability by representing nonlinear transformations as learnable univariate functions, enabling direct analysis of feature-response relationships. However, functional interpretability alone may not fully characterize models with structured latent spaces. HYDRA complements functional analysis by examining hyperbolic representation dynamics to provide additional insights into model behavior.
Method
Problem Setup and Architecture: We consider supervised learning with vector-valued inputs D = (xi, yi) ni=1, where xi ∈ Rp is a normalized input and yi is either a continuous target or a binary label. The goal is to learn a predictor fθ(x) that is accurate, parameter-efficient, and inspectable. Dense KAN layers replace scalar weights with learnable univariate functions, but a hidden-to-hidden KAN block assigns one function to each coordinate pair. This creates a dominant O(d2K) spline cost for hidden width d and K spline bases.
HYDRA addresses this cost by combining three operations. First, it represents hidden states in a bounded Poincaré ball, so latent scale can be encoded through radius as well as direction. Second, it performs KAN-style functional learning in the tangent space, where ordinary one-dimensional spline bases remain available. Third, it compresses the tangent update through a learned prototype space before mapping the state back to the hyperbolic manifold.
For layer l, HYDRA maps the current hyperbolic state to the tangent space, applies a low-rank spline update, and reconstructs a bounded hyperbolic state:
zl = logc0(hl),
pl = W↓ zl,
sl = Φl(pl),
z̃l+1 = zl + αl W↑ sl,
hl+1 = Πrl+1(expc0(z̃l+1)),
where Φl is a spline block in prototype coordinates, W↓ and W↑ define the bottleneck, αl is a learned residual scale, and Πrl+1 enforces the layer radius budget. This formulation keeps local function learning Euclidean while making the hidden trajectory geometrically constrained and measurable.
Hyperbolic Embedding and Tangent-Space Updates: HYDRA represents hidden states on the Poincaré ball Bcd = h ∈ Rd: c∥h∥22 0. The normalized input is first mapped to a tangent vector u0 = gemb(x). The initial hidden state is then formed by a bounded exponential map h0 = expc0(ũ0), with ∥h0∥c ≤ remb, where remb is the embedding radius budget. The embedding map gemb is linear by default and can be replaced by a spline-based input KAN when stronger input response functions are needed.
The exponential and logarithmic maps at the origin are:
expc0(v) = tanh(√c∥v∥2) v/(√c∥v∥2),
logc0(h) = artanh(√c∥h∥2) h/(√c∥h∥2).
They are interpreted by continuity at the origin. These maps define a fixed coordinate chart for the KAN update. The model therefore avoids designing coordinate-wise spline functions directly on the manifold, while still returning to a hyperbolic state after each block.
Given zl = logc0(hl), a tangent-space KAN update has the residual form zl+1 = zl + αl Δl(zl). For a dense KAN layer, the ith coordinate is Δi(z) = Σj=1d ϕij(zj), with ϕij(t) = Σm=1K aijm Bm(t), where Bm denotes spline bases and aijm denotes learned coefficients. This form is expressive and interpretable, as each edge has an explicit response curve, leading to every input-output coordinate pair owning a separate spline. HYDRA keeps the same spline principle but moves the costly functional operator into a lower-dimensional prototype space.
Low-Rank Prototype Functional Learning: The low-rank block starts from the observation that hyperbolic radius can carry part of the latent scale variation. The tangent update need not always use all d hidden directions independently. HYDRA therefore learns an r-dimensional prototype representation, applies the KAN operator there, and lifts the result back to the hidden dimension:
pl = W↓ zl,
Δl(zl) = W↑ Φl(pl),
where W↓ ∈ Rr×d, W↑ ∈ Rd×r, r ≪ d.
The spline block Φl is applied before the up-projection, so the update is a nonlinear functional operator whose learned spline interactions are shared through the prototype coordinates. A full hidden-to-hidden KAN block contains approximately Pfull ≈ d2(K+1) parameters, whereas the prototype block uses Plr = 2dr + r2(K+1). The approximate compression ratio is Plr/Pfull ≈ (2r/(K+1)d) + (r/d)2. The low-rank prototype design reduces the functional parameter cost from O(d2K) to O(dr + r2K), where the compression ratio is mainly determined by the rank ratio r/d. Thus the dominant spline cost depends on r rather than d. A small rank is useful when the tangent update has low effective dimension, while a larger rank can be selected when the task requires more independent interactions.
Radius Control Mechanism: Hyperbolic representations can become unstable near the boundary of the Poincaré ball. In that region, distances and tangent coordinates are amplified, so an unconstrained model may reduce the loss by pushing samples outward instead of learning smoother tangent-space functions. HYDRA controls this behavior with two mechanisms. The hard projection Πrl keeps every layer within a prescribed radius. A soft penalty discourages unnecessary outward movement before the projection becomes active:
Lrad = (1/(L+1)) Σl=0L q [max(0, ρ(hl) − rallow,l)].
Here ρ(hl) is the normalized radial coordinate, rallow,l is the soft threshold, and q controls the penalty shape. In practice, rallow,l = τ rl with 0 < τ ≤ 1, where rl is the hard radius budget. Projection and penalty are complementary: projection prevents boundary violations, while the penalty changes the training direction before the representation reaches the high-amplification region.
After the final HYDRA block, prediction is made from the tangent coordinate zL = logc0(hL), with ŷ = gout(zL). The readout gout is linear by default. For regression, HYDRA minimizes mean squared error on normalized targets. For binary classification, it minimizes binary cross-entropy with logits. The complete objective is L = Lsup + λrad Lrad + λsp Lsp, where Lsup is the supervised loss and Lsp combines sparsity and smoothness regularization on the spline coefficients. This objective makes HYDRA a radius-constrained, low-rank functional model.
Experiments and Results
Experimental Setup: We evaluate HYDRA on eight widely used tabular benchmark datasets collected from the OpenML platform (Vanschoren et al. 2013): CCPP, Energy Heating, Parkinsons Telemonitoring, Real Estate Valuation, Heart Statlog, Ionosphere, Phoneme, and QSAR Biodegradation. These datasets cover both regression and binary classification tasks and exhibit diverse characteristics in terms of sample size, feature dimensionality, and underlying data distributions. The first four datasets are formulated as regression problems, whereas the remaining four are treated as binary classification tasks. Regression performance is evaluated using RMSE, whereas classification performance is measured using accuracy. For all experiments, the dataset is randomly split into training, validation, and test sets with ratios of 80%, 10%, and 10%, respectively, using a fixed random seed of 42.
Implementation details: All experiments are conducted on a single NVIDIA L20 GPU, and all models use the same preprocessing pipeline, optimizer configuration, and evaluation protocol. Regression models predict a single scalar and are evaluated on the original target scale, while classification models output a single logit that is converted into a probability for evaluation. Beyond the task-specific loss, HYDRA introduces only two additional regularization terms: a spline regularizer that encourages smooth and compact spline functions, and a radius regularizer that constrains hidden representations within a layer-wise radius budget. Therefore, the observed performance differences mainly reflect architectural improvements rather than variations in optimization strategies, training procedures, or computational resources.
Evaluation Criteria: We select models by the primary metric and report trainable parameter count to measure compactness. HYDRA hyperparameters include hidden width, prototype rank r, spline resolution, learning rate, weight decay, radius budget, and regularization strength. The rank is treated as a compression knob: we choose the smallest rank that preserves the main metric when possible. Parameter-free geometric operations, including exponential maps, logarithmic maps, and radius projection, are not counted as trainable capacity. The parameter-efficiency claim is evaluated primarily against Euclidean KAN and MLP because they share the same dense functional or dense hidden-state modeling role; the remaining baselines are included to contextualize predictive accuracy and interpretability-oriented alternatives.
The selection rule separates predictive quality from compactness. First, we search for HYDRA configurations that are competitive on the primary metric. Second, among configurations with similar primary performance, we prefer the one with fewer trainable parameters. This avoids two misleading extremes: a large HYDRA variant that wins mainly by capacity, and a very small HYDRA variant whose parameter count is attractive but whose accuracy is no longer competitive. The reported configuration is therefore a parameter-performance trade-off rather than the largest model found during tuning.
Main Benchmark Results: Table 1 shows that HYDRA achieves the strongest or tied-strongest primary metric on all eight datasets. On the regression tasks, HYDRA improves over KAN and MLP on CCPP, Energy Heating, Parkinsons Telemonitoring, and Real Estate Valuation while using fewer parameters than both Euclidean counterparts. On the classification tasks, HYDRA matches the best accuracy on Heart Statlog and achieves the highest accuracy on Ionosphere, Phoneme, and QSAR Biodegradation. For example, on the Parkinsons Telemonitoring dataset, HYDRA reduces RMSE from 4.424 of KAN to 3.534 while decreasing trainable parameters from 2.4k to 1.4k, corresponding to a 20.1% performance improvement with 41.7% fewer parameters. Compared with MLP, HYDRA further reduces RMSE by 33.4% under the same parameter reduction. These results support the central claim that HYDRA is not merely accurate, but accurate under a smaller trainable budget.
The benchmark mixes regression and classification because the two settings stress different aspects of the model. Regression datasets such as CCPP and Real Estate require smooth response surfaces without allocating a large number of spline edges. Classification datasets such as Ionosphere, Phoneme, and QSAR require separation while discouraging uncontrolled movement toward the Poincaré boundary. HYDRA performs well in both regimes, suggesting that the hyperbolic coordinate is not only acting as a classifier margin and not only as a regression smoother. Instead, it provides an internal scale coordinate that can be coupled with local spline response across objectives.
Parameter Efficiency: The hyperbolic maps themselves are parameter-free under fixed curvature, so HYDRA's parameter savings must come from smaller hidden widths or lower prototype ranks. Compared with KAN and MLP, HYDRA uses fewer parameters on every dataset. This matches the low-rank analysis: if the radial coordinate absorbs part of the scale-like variation, fewer prototype directions are needed for the KAN update. It is useful to distinguish three sources of trainable capacity. The first is the input embedding, which maps raw features into the hidden representation. The second is the hidden-to-hidden functional update. This is the dominant term for dense KAN-style models because each hidden input-output pair can own a separate spline. The third is the output readout, which is small for the widths used here. HYDRA targets the second term: the down projection, prototype spline block, and up projection replace a dense functional map with a compact bottleneck. Thus the parameter reduction is architectural rather than a post-hoc pruning effect. When the tangent update has low effective dimension, a small rank preserves the relevant directions; when many independent interactions are needed, the selected rank grows and the savings weaken.
Ablation Studies: The low-rank ablation asks whether prototype compression preserves performance relative to a full-rank HYDRA block. The radius ablation asks whether the hyperbolic radius budget improves optimization beyond numerical safeguarding. The selected low-rank models use a mean of 46.8% and a median of 33.8% of the corresponding full-rank HYDRA parameters. Compression is strongest on Heart Statlog, CCPP, Phoneme, and QSAR, while Real Estate and Ionosphere require ratios close to one, showing that the benefit is task-dependent. Across all eight datasets, the constrained model has a smaller mean radius than the unconstrained model, and the primary metric improves in the selected comparisons. Radius control pulls representations inward while preserving or improving the primary metric. This matches the theoretical role of the radius budget: it bounds distance amplification and log-map gradients, discouraging near-boundary shortcuts. Without radius control, the model may exploit the outer region of the Poincaré ball, where even small Euclidean perturbations correspond to large hyperbolic distances. Such movement can help separate samples, but it can also create an unstable shortcut in which the model reduces the loss by pushing representations outward instead of learning a smooth tangent-space functional response. The constrained variant limits this behavior. The simultaneous reduction in mean radius and improvement in the primary metric suggests that radius control changes the learned representation, not only the numerical range of the hidden state.
Interpretability Analysis: HYDRA's interpretability comes from the geometry of its hidden representation. For a controlled feature sweep, we record the final hyperbolic radius, trajectory shape, and path length of the hidden state. These quantities describe how the model reorganizes its internal representation as one input variable is changed. At the same time, we have introduced SHAP (SHapley Additive exPlanations) values as a reference to demonstrate the interpretability of the model. SHAP is a game-theoretic interpretability method that quantifies the marginal contribution of each input feature to a model prediction. By decomposing predictions into feature-level positive or negative effects, SHAP provides an intuitive explanation of complex model behavior.
For a sweep t ↦ (x−j, t) with final representation γj(t), the path length Lj = ∫ ∥γj′(t)∥Bc dt, with ∥γj′(t)∥Bc = λcγj(t) ∥γj′(t)∥2, weights Euclidean displacement by the local conformal factor. Large radius changes or long paths therefore indicate stronger internal reorganization.
We used the CCPP dataset as a case study to examine whether HYDRA's latent geometry provides physically meaningful interpretability. The prediction target is net hourly electrical power output. Ambient temperature (AT), ambient pressure (AP), and relative humidity (RH) mainly describe gas-turbine operating conditions, whereas exhaust vacuum (V) reflects the steam-turbine side of the combined-cycle process. The AT sweep produced the clearest radius-output relation. The trajectory moved from a small-radius, positive-SHAP regime at low AT to a large-radius, negative-SHAP regime at high AT, and the one-feature PDP reduced predicted power by 37.36 MW. This agrees with the physical expectation that hotter intake air lowers air density and reduces the available mass flow through the gas turbine. The V sweep also showed a negative output direction, but with a longer and more tortuous hyperbolic path. This pattern is consistent with a coupled steam-side condition that changes the internal representation, rather than acting as a simple direct driver. The corresponding AP and RH sweeps induced weaker geometric responses, consistent with their secondary roles among the operating variables. Overall, these observations support hyperbolic radius and path geometry as HYDRA-specific interpretability diagnostics.
Conclusion
We introduced HYDRA as a compact extension of KAN that couples tangent-space spline computation with hyperbolic latent representations. Across eight tabular benchmarks, HYDRA achieved competitive or superior predictive performance while reducing trainable parameters by 34.9% relative to Euclidean KAN and by 37.1% relative to MLP on average. Ablation studies showed that the low-rank prototype block retained predictive performance, and that radius control reduced near-boundary saturation in the Poincaré ball. The CCPP interpretability case study further showed that hyperbolic radius and latent trajectories can provide HYDRA-specific diagnostic signals that are consistent with physical expectations and SHAP trends. Overall, these results indicate that hyperbolic representation geometry can support parameter-efficient KAN-style function learning while offering an inspectable latent-space view of model behavior.
Improvements for AI systems
Improvements to AI Systems:
- Parameter-Efficient Functional Learning with Low-Rank Prototype Bottlenecks
-
Replace dense per-edge spline functions in KANs with a low-rank prototype block (down-projection → spline block → up-projection). This reduces hidden-to-hidden parameter complexity from O(d2K) to O(dr + r2K), enabling larger hidden widths without proportional parameter growth.
-
The improved system can scale to high-dimensional tabular or sensor data with limited memory, achieving accuracy comparable to full-rank KANs while using 35–47% fewer parameters.
- Hyperbolic Latent Space with Radius Control for Stable Representation Learning
-
Map hidden states into a bounded Poincaré ball and perform tangent-space updates, with explicit hard projection and soft radius penalties to prevent near-boundary saturation.
-
The improved system can learn compact, interpretable representations for data with latent magnitude variations (e.g., energy output, medical telemetry) without instability or shortcut learning, improving generalization and training robustness.
- Interpretable Latent Trajectory Diagnostics
-
Use hyperbolic radius, path length, and trajectory shape as model-specific interpretability signals, complementing SHAP values.
-
The improved system can provide physically meaningful explanations (e.g., temperature effects on power output) by showing how input changes reorganize internal representations, aiding domain experts in validating model behavior.
- Task-Adaptive Compression via Rank Selection
-
Treat prototype rank as a tunable compression knob, selecting the smallest rank that preserves primary metric performance.
-
The improved system can automatically balance accuracy and parameter efficiency across diverse tasks (regression vs. classification), avoiding over-parameterization while maintaining competitive predictive power.
- Unified Handling of Regression and Classification with Geometric Regularization
-
Combine supervised loss with spline smoothness and radius penalties in a single objective, using the same architecture for both continuous and binary targets.
-
The improved system can generalize across mixed task types (e.g., predicting continuous outputs like energy load and classifying anomalies) without task-specific architectural changes, reducing engineering overhead.
- Bounded Input Embedding for Stable Initialization
-
Map normalized inputs to a bounded hyperbolic embedding with a fixed radius budget, preventing extreme initial states.
-
The improved system can train more stably from scratch, especially with deep or wide architectures, by avoiding large initial tangent-space gradients.
What the Improved AI System Can Do:
-
Learn compact, accurate models for tabular data (e.g., energy forecasting, medical diagnostics, material property prediction) with up to 40% fewer parameters than standard KANs or MLPs.
-
Provide interpretable, geometry-based explanations of feature effects (e.g., which input drives output changes and how strongly), useful for scientific discovery and model auditing.
-
Operate reliably on resource-constrained devices (e.g., edge sensors) due to reduced memory footprint, while maintaining state-of-the-art accuracy.
-
Adapt compression automatically per task, ensuring no loss in predictive quality when parameter budget is tight.
Sources
- Wav-KAN: Wavelet Kolmogorov-Arnold Networks
- Kolmogorov-Arnold Networks are Radial Basis Function Networks
- KAN 2.0: Kolmogorov-Arnold Networks Meet Science
- Chebyshev Polynomial-Based Kolmogorov-Arnold Networks: An Efficient Architecture for Nonlinear Function Approximation
- PRKAN: Parameter-Reduced Kolmogorov-Arnold Networks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks