TESLA: Taylor Expansion of Sinusoidal Learnable Activations
Daehwa Ko, Jaehyeon Kim, Seunghyun Ham, Jay Hoon Jung
Korea Aerospace University
cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 15 pages, 7 figures. Accepted at AISTATS 2026
Journal ref: Proceedings of the 29th International Conference on Artificial Intelligence and Statistics (AISTATS 2026), PMLR 300, 2026
Code: https://github.com/KAU-QuantumAILab/TESLA
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: TESLA: Taylor Expansion of Sinusoidal Learnable Activations Authors: Daehwa Ko, Jaehyeon Kim, Seunghyun Ham (Korea Aerospace University), Jay Hoon Jung Abstract: The parity problem—deciding whether
Terminology
Summary
TESLA: Taylor Expansion of Sinusoidal Learnable Activations
Abstract:
The parity problem—deciding whether the number of ones in a binary vector is odd or even—remains challenging for standard neural networks due to linear inseparability and the need for global interactions. We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components. Theoretically, we show that constraining TESLA’s coefficients yields Lipschitz/Rademacher complexity bounds and shapes the training dynamics to emphasize higher-frequency structure. Empirically, on parity with input length n = 32, TESLA attains strong generalization with 100K training samples (≈ 0.002% of the 2 32 input space) and remains robust under heavy corruption, retaining high accuracy with up to 30% label noise. We also compare against periodic and frequency-based baselines (SIREN, SNAKE, and Fourier feature embeddings) on parity and Forrelation. Beyond synthetic structure, TESLA delivers comparable performance on ImageNet-100, indicating that activation-level degree control transfers to more general vision workloads.
1. Introduction
Modern neural networks largely build complex functions by stacking local nonlinearities such as ReLU and GeLU with linear maps. These designs excel at capturing local structure in images and language, but they can be inefficient for global or long-range interaction tasks that depend on high-order combinations of many input coordinates, global parity/phase, or other global symmetries. Empirically and theoretically, traditional activation functions and vanilla MLPs exhibit a spectral bias toward low-frequency (low-degree) components, so representing strong high-order or global structure often requires substantially greater depth, width, or very large weight norms.
To enable direct and learnable control over the polynomial order at the activation level, the authors propose a simple parametric activation function based on a finite sine and cosine basis. This formulation allows the network to learn Fourier-like coefficients that explicitly control the effective polynomial order of activations. They refer to this activation as Taylor Expansion of Sinusoidal Learnable Activation (TESLA).
Main contributions:
-
They derive tight Lipschitz bounds, establish Rademacher complexity–based generalization guarantees, and characterize learning dynamics, showing that TESLA’s inductive bias naturally favors recovery of low- and mid-frequency structure.
-
They present constructive approximation results for ridge- and interaction-type functions, together with complementary lower bounds showing that standard piecewise-linear activations require substantially more resources to match TESLA’s ability to represent oscillatory structure.
-
On synthetic benchmarks, TESLA generalizes a 32-bit parity task with 100K training samples and remains robust under label noise up to 30%. It also improves performance on Forrelation and LPN tasks and remains usable in ImageNet-100-scale models while incurring negligible throughput overhead (≈ 1%).
2. Related Work
Recent research has explored various ways to enrich neural network representations through activation design. Several works, such as PReLU and APL, introduce parametric nonlinear activations with a small number of learnable parameters to improve expressiveness and facilitate optimization. While these methods directly modify the pointwise nonlinearity, they do not provide explicit mechanisms for controlling polynomial or spectral degree. In contrast, approaches that learn per-edge or per-unit activations—such as KANs and spline-based methods—offer substantially greater expressivity and can closely approximate complex mathematical or scientific functions. However, prior studies and surveys have noted that this increased flexibility can lead to overfitting or instability in low-data regimes unless it is carefully regulated through techniques such as regularization or parameter sharing.
A complementary line of work employs sinusoidal and Fourier-based parameterizations to address the spectral limitations of standard activations and embeddings. Periodic activations (e.g., SIREN) and positional or Fourier feature encodings mitigate spectral bias by introducing high-frequency components into the network—either by expanding its effective spectral support or by modifying the spectrum of the initial Neural Tangent Kernel (NTK). The NTK is a theoretical construct that characterizes the training dynamics of infinitely wide neural networks, showing that such networks evolve like linear models governed by a fixed kernel, which in turn guarantees convergence to a global minimum.
In parallel, a substantial body of theoretical work has emerged to analyze spectral bias (or the frequency principle) and NTK dynamics, aiming to explain why standard networks prioritize low-frequency components and how modifying kernels or embeddings can alter this behavior. Classical approximation-theoretic results—such as Barron-type bounds, depth–width separation theorems, and lower bounds for piecewise-linear approximations—provide a rigorous framework for comparing different parameterizations.
3. Proposed Method
3.1 Activation Function Definition
The authors define a parametric activation function, shared across all neurons within a layer, built from a finite sine and cosine basis:
ϕK(z) = Σ k=1 K [(a k/k) sin(kz) + (b k/k) cos(kz)]
where K ∈ N controls the maximum frequency, and a k, b k k=1 K are learnable scalar coefficients. For a neuron receiving input z = v T x, the neuron output is f(x) = ϕK(v T x). They also define the coefficient budget:
A K:= Σ k=1 K (a k + b k)
which appears in stability and complexity bounds.
3.2 Taylor (Maclaurin) Expansion and Polynomial Coefficients
Differentiation of Eq. (1) gives: ϕ′K(z) = Σ k=1 K [a k cos(kz) − b k sin(kz)].
Expanding ϕK via the Taylor series of sine and cosine gives an expression showing that each polynomial order is an explicit linear functional of the learned coefficients a k, b k, giving direct control over degree components at the activation level.
Interpretation vs. Implementation: TESLA is implemented as the finite trigonometric expansion in Eq. (1); the authors do not evaluate a polynomial series during the forward pass. They use Eq. (3) as an interpretation that links learned trigonometric coefficients to effective polynomial degree and motivates the coefficient budget in their analysis.
Input Spectralization vs. Activation Spectralization: Input-side spectralization (e.g., Fourier features) maps inputs to high-frequency coordinates and leaves mode selection to later layers. TESLA instead spectralizes the activation itself: coefficients a k, b k directly shape Maclaurin-order terms, enabling selective amplification or attenuation of target orders. In practice, this design reduces the need for handcrafted frequency mappings at the input stage and moves spectral control into a small set of learnable activation coefficients.
4. Theoretical Analysis
4.1 Derivative Bound and Stability
Since sin(·) ≤ 1 and cos(·) ≤ 1, each term is bounded by its amplitude, a k cos(kz) − b k sin(kz) ≤ sqrt(a k2 + b k2); therefore, ∥ϕ′K∥∞ ≤ Σ k=1 K sqrt(a k2 + b k2) ≤ Σ k=1 K (a k + b k) = A K.
For an L-layer network f(x) = W L ϕK(· · · ϕK(W 1 x)), the Lipschitz constant satisfies Lip(f) ≤ (Π l=1 L ∥W l∥ op) A K L−1.
For an L loss-Lipschitz loss L, ∥∇ W l L(f(x), y)∥ F ≤ L loss ∥x∥ 2 (Π j=1 L ∥W j∥ op) A K L−1. This highlights the trade-off: increasing K expands representable high frequencies, while controlling A K keeps optimization stable.
4.2 Rademacher Complexity and Generalization
Theorem 1: Assume ∥x i∥ 2 ≤ R for all i = 1,..., N. Define F K = x ↦ ϕK(v T x) ∥v∥ 2 ≤ W, A K ≤ A, and B 0(A):= sup ϕK(0) for ϕK with A K ≤ A. Then, for any sample of size N, the empirical Rademacher complexity satisfies:
R̂ N(F K) ≤ (AW R + B 0(A)) / √N
Consequently, for any L loss-Lipschitz loss L, the generalization gap scales as O(L loss (AW R + B 0(A)) / √N). The bound recovers the familiar linear-class rate Θ(W R / √N) up to the multiplicative factor A due to the activation Lipschitz constant.
4.3 Mode-Wise Learning Dynamics
The authors analyze learning dynamics on the circle T = [0, 2π) in the Fourier basis. Under squared loss, the linearized gradient flow around initialization is ∂τ(fτ − f⋆) = −ηK ϕ(fτ − f⋆), where K ϕ is the kernel operator.
With TESLA coefficients a k, b k k=1 K, the Jacobian features satisfy ∂fθ/∂a k = (1/k) sin(kt) and ∂fθ/∂b k = (1/k) cos(kt), so K ϕ is translation-invariant with kernel K ϕ(t, t′) = Σ k=1 K (1/k2) cos k(t − t′).
The orthonormal Fourier basis ψ m,s(t) = √2 sin(mt) and ψ m,c(t) = √2 cos(mt) diagonalizes K ϕ: K ϕ[ψ m,a] = λ m(ϕ) ψ m,a, with λ m(ϕ) = 1/(2m2) for a ∈ s, c, 1 ≤ m ≤ K, and λ m(ϕ) = 0 for m > K.
The dynamics decouple as e m,a(τ) = e m,a(0) exp −ηλ m(ϕ) τ, and ∥fτ − f⋆∥2 L2 = Σ m,a e m,a(τ)2. Thus, larger λ m(ϕ) yields faster decay of the corresponding mode. Analytically, TESLA gives λ m(ϕ) ∝ m−2; empirically, its spectrum is flatter than RF proxies, allocating relatively larger eigenvalues to medium–high modes and accelerating recovery of higher-order structure.
4.4 Theory-Guided Choice of the Harmonic Count K
The authors give a compact, operational rule for how the harmonic count K should scale with task complexity and sample size by balancing approximation and estimation.
For ε ∈ (0, 1/2), define the effective interaction order m eff:= inf m ∈ 0,..., d: Σ T≤m f̂ T2 ≥ 1 − ε. m eff is the interaction order needed to capture most of the target energy (e.g., parity of order m has m eff = m).
Assuming the best-in-class L2 approximation error over F K satisfies inf E((f(x) − f⋆(x))2) ≤ B(K/m eff), where B(z) ∈ c1 e−γz, c1 z−p, and the estimation term satisfies R̂ N(F K) ≤ C 0 (A K / √N) √log(2K), the excess-risk bound is:
E(f̂ K) ≲ B(K/m eff) + C (A K / √N) √log(2K) + C √(log(1/δ)/N)
Balancing the first two terms gives K⋆ = Θ(m eff · κ(N, A K)), where κ varies slowly with N and A K: for B(z) = c1 e−γz, κ = Θ(log N); for B(z) = c1 z−p, κ = Θ(N 1/(2p) (log N) 1/(2p)) up to constants. Intuitively, taking K much larger than m eff dilutes the per-harmonic budget (average amplitude ∼ A K/K) and increases estimation error.
In d-bit Parity, m eff = d, hence the theory predicts K⋆ = Θ(d) up to slowly varying factors. Empirically, they find a stable range κ ∈ [0.2, 0.4] at N = 10 5, placing the optimum near K ≈ d/4.
5. Experiments
5.1 Comparisons with Periodic and Frequency-based Baselines
To provide a balanced comparison with periodic and Fourier-style activations, the authors include SIREN, SNAKE, and a Fourier-feature embedding baseline in both parity and Forrelation experiments, using matched architectures and training protocols unless otherwise noted. For the Fourier-feature embedding baseline, they use a 64-dimensional input mapping with sin(2 k πx) and cos(2 k πx) for k = 0:31, followed by the same MLP architecture with ReLU.
5.2 Continuous-domain Evaluation
The authors evaluate TESLA on three representative settings: (i) physics-informed neural networks (PINNs) for PDE solving, (ii) implicit neural representations (INRs) for coordinate-based image reconstruction, and (iii) mixed-frequency regression for explicit high-frequency signal modeling.
Physics-informed neural networks (PINNs): They train a PINN for the 1D viscous Burgers’ equation. The total loss is L total = L PDE + 100 L IC + 100 L BC. As shown in Table 1, TESLA achieves the lowest final loss on the Burgers’ equation PINN among the compared activations (L total = 1.21 × 10−3 vs. 7.13 × 10−1 for ReLU, 1.19 × 10−2 for Tanh, 1.04 × 10−2 for SIREN).
Implicit neural representations (INRs): They evaluate INRs on Kodak24 and DIV2K, where a 6-layer MLP (512 hidden units) is trained per image for 3000 epochs under an identical training budget. On Kodak24, TESLA achieves PSNR 27.72 and SSIM 0.9671, compared to SIREN's 25.33 PSNR and 0.9404 SSIM. On DIV2K, TESLA achieves PSNR 24.73 and SSIM 0.9522, compared to SIREN's 22.77 PSNR and 0.9301 SSIM.
Mixed-frequency regression: They regress a target signal composed of mixed sine/cosine components from 3.29 Hz to 79.90 Hz. TESLA achieves test MSE of 0.0100, the only method below 0.011 (ReLU: 0.0681, GeLU: 0.0756, SiLU: 0.0756, SIREN: 0.0248).
5.3 Parity with Label Noise: Task, Metrics, and Interpretation
Task: They evaluate on Parity and global-statistics problems where the target depends on a high-order global interaction of the input bits. Let x ∈ 0, 1 d and S ⊆ [d] with S = m. The clean label is y clean = ⊕ i∈S x i ∈ 0, 1, i.e., the parity of the selected bits. This benchmark is intentionally adversarial to local, piecewise-linear activations, as the Bayes-optimal classifier is a global rule with a single nonzero Fourier coefficient at frequency S.
Data Generation, Splits, and Noise Settings: For each run, they sample i.i.d. inputs uniformly from 0, 1 d and assign splits deterministically via a fixed hash partition h(x) mod 3, ensuring disjoint and reproducible train/validation/test sets (i.e., no leakage). They then sample without replacement within each split, using 100,000 training samples and 20,000 validation samples for all d. They evaluate three settings: (i) clean parity, (ii) noisy parity with independent label flips p ∈ 0.05, 0.10, 0.20, 0.30, and (iii) LPN; in the noisy setting, train and test use the same flip rate p.
LPN: They also evaluate the classical LPN (Learning Parity with Noise) variant, where the relevant subset is unknown and must be inferred from data. Fix a hidden secret vector s ∈ 0, 1 d and draw query vectors a ∼ Unif(0, 1 d) i.i.d. The observed label is y = ⟨a, s⟩ mod 2 ⊕ e, e ∼ Bernoulli(η).
Metrics and Bayes Limit: With symmetric label noise (flip rate p < 1/2) on the test set, perfect recovery of the underlying rule cannot yield 100% measured accuracy. The Bayes-optimal expected test accuracy under symmetric, input-independent label flips at rate p is A max(p) = 1 − p.
5.4 Forrelation
Forrelation is a decision problem on pairs of Boolean functions f, g: 0, 1 n → ±1. Let M = 2 n. The goal is to decide whether the correlation Φ f,g = (1/√M) ⟨Hf, g⟩ is high or low under a standard promise. Here H is the Hadamard matrix and ⟨·, ·⟩ denotes the inner product. They treat Φ f,g ≥ 0.6 as positive and Φ f,g ≤ 0.01 as negative. The property of being forrelated is global. A classifier must integrate information spread across the entire pair of truth tables. They use a two-layer MLP that takes as input the concatenation of the truth tables of f and g with length 2 n+1 and outputs a single logit for binary prediction. They train and evaluate on a synthetic dataset with n = 12, consisting of 10,000 training examples and 10,000 test examples.
5.5 ImageNet-100 Classification
To empirically validate the performance and efficiency of TESLA, they conducted experiments on the ImageNet-100 dataset. To demonstrate TESLA’s general applicability, their evaluation includes both Transformer-based models (ViT, MLP-Mixer) and CNN-based models (ResNet, MobileNetV3). They compare TESLA’s performance against conventional activation functions, namely ReLU, GeLU, and SiLU. TESLA uses K = 2 for ViT-T/16, ResNet-18, and MobileNetV3, and K = 6 for MLP-Mixer. The evaluation is based on a comprehensive analysis of not only the classification performance (Top-1 Accuracy) but also key computational efficiency metrics, including FLOPs, throughput (imgs/s), and the number of parameters.
Results from Table 4:
-
ViT-T/16: TESLA achieves 62.53% Top-1 (vs. 62.49% ReLU, 63.67% GeLU, 63.60% SiLU), with FLOPs 1.09 GFLOPs and throughput 7,509 imgs/s.
-
MLP-Mixer-b16: TESLA achieves 58.79% Top-1 (best among compared), with FLOPs 12.75 GFLOPs and throughput 1,139 imgs/s.
-
ResNet-18: TESLA achieves 74.01% Top-1 (best among compared), with FLOPs 1.82 GFLOPs and throughput 8,433 imgs/s.
-
MobileNetV3-S: TESLA achieves 67.82% Top-1, with FLOPs 0.05 GFLOPs and throughput 35,483 imgs/s.
5.6 Result Analysis
Parity: Figure 3 shows the performance of TESLA and baseline models. For the parity experiments they use a fully-connected MLP with 2 hidden layers and hidden dimension H = 128. Across the plotted settings, TESLA matches the Bayes limit across noise levels while remaining near 100% on the clean task. For p ∈ 0.05, 0.10, 0.20, 0.30, TESLA’s final accuracy concentrates near 95%, 90%, 80%, 70%. In contrast, common activations always collapse toward 50% as the number of bits grows, which is no better than random guessing of parity. In the LPN setting, TESLA sustains high accuracy up to 27-bits before degrading in the hardest regimes.
Periodic baselines on parity: Table 5 compares TESLA against periodic and frequency-based baselines. For easier settings (16–20 bits), TESLA and SIREN both reach 100% accuracy under no noise. However, as the bit-length and noise increase, SIREN collapses to random guessing (≈ 50%) while TESLA remains stable (e.g., at 24 bits with 30% noise, TESLA achieves 65.53% vs. 49.47% for SIREN; at 32 bits with no noise, TESLA achieves 95.87% vs. 50.10% for SIREN). SNAKE remains near chance across all settings, and Fourier-Emb similarly fails to capture the global interaction signal in these regimes.
Forrelation: Figure 4 reports validation accuracy over epochs for hidden sizes H = 128, 512, 1024. TESLA consistently learns faster and reaches higher final accuracy than standard activations across all K settings. Larger hidden sizes improve performance for every method. The gap in favor of TESLA remains, which indicates that TESLA captures the global correlation signal more effectively.
ImageNet-100: Table 4 reports results on ViT-T/16, MLP-Mixer, ResNet-18, and MobileNetV3. To ensure architecture-aware comparison, they keep activations inside convolutional blocks unchanged for CNNs and replace only non-convolutional activations (e.g., MLP heads, projection/FFN layers) with TESLA; for ViT and MLP-Mixer, they replace all pointwise activations. TESLA is trained with an explicit l1 penalty enforcing a layer-wise coefficient budget A K for stability and generalization control. Across models, TESLA remains competitive in Top-1 accuracy while keeping FLOPs and throughput within about 1% of standard activations.
Stability vs. sinusoidal baselines: Table 6 compares TESLA against SIREN under the same ImageNet-100 training settings. In these runs, SIREN exhibits severe optimization instability and much lower accuracy (e.g., ViT-T/16: 4.11% for SIREN vs. 62.53% for TESLA; ResNet-18: 37.41% for SIREN vs. 74.01% for TESLA), whereas TESLA trains stably and reaches standard accuracy levels.
6. Future Work
TESLA opens several promising avenues for future research. Theoretically, the authors plan to extend their continuous Fourier and NTK analysis to discrete Boolean domains to establish rigorous approximation bounds and explain the empirical gains observed on parity-type tasks. On the systems side, they will investigate sparse factorizations, low-bit quantization, and hardware-aware implementations to preserve TESLA’s spectral benefits while minimizing inference latency. Finally, they aim to integrate TESLA into large-scale Transformers and LLMs—such as augmenting FFN activations or positional encodings—to evaluate its impact on long-range reasoning, compositional generalization, and sample efficiency.
7. Conclusion
In this work, the authors introduced TESLA, a learnable sinusoidal activation that enables explicit degree/frequency control while preserving stable optimization through coefficient budgeting. They provided Lipschitz and Rademacher-complexity bounds and analyzed mode-wise learning dynamics to connect the parameterization to frequency-selective behavior.
Empirically, TESLA consistently improves tasks that require global, high-order interactions. On parity with label noise, it remains well above chance at larger bit-lengths and heavy corruption, and is more stable than periodic/frequency-based baselines (SIREN, SNAKE, and Fourier feature embeddings) in the hardest regimes. On Forrelation, TESLA achieves the highest accuracy across widths, while periodic and Fourier-style baselines remain near chance. On continuous-domain tasks (PINNs, INRs, and mixed-frequency regression), TESLA is competitive with or better than both standard and sinusoidal activations. On ImageNet-100, TESLA remains competitive across architectures with modest overhead and is substantially more stable than SIREN. Overall, TESLA suggests that activation-level degree control is a robust, practical alternative to input-only spectralization.
Improvements for AI systems
Based on the paper, here are specific improvements to AI systems and what the improved systems can do:
1. Replace standard activations (ReLU, GeLU, SiLU) with TESLA in MLP layers for tasks requiring global, high-order feature interactions
-
The improved system can solve parity problems up to 32 bits with only 100K training samples (0.002% of input space), whereas standard activations collapse to random guessing beyond 16 bits.
-
It maintains 70% accuracy even with 30% label noise on parity tasks, versus 50% (chance) for standard activations.
2. Use TESLA instead of SIREN or Fourier feature embeddings for implicit neural representations (INRs) and physics-informed neural networks (PINNs)
-
The improved system achieves PSNR 27.72 on Kodak24 image reconstruction (vs. 25.33 for SIREN) and SSIM 0.9671 (vs. 0.9404).
-
For PINNs solving Burgers' equation, it reduces total loss to 1.21×10−3, three orders of magnitude better than ReLU (7.13×10−1) and 10× better than Tanh/SIREN.
3. Apply TESLA to mixed-frequency regression tasks where signals contain multiple high-frequency components
- The improved system achieves test MSE of 0.0100 on signals spanning 3.29 Hz to 79.90 Hz, outperforming SIREN (0.0248) and standard activations (0.068–0.076).
4. Integrate TESLA into Transformer and CNN architectures for vision tasks with minimal overhead
-
The improved system achieves 74.01% Top-1 accuracy on ImageNet-100 with ResNet-18 (best among ReLU, GeLU, SiLU) while maintaining throughput within 1% of standard activations.
-
For MLP-Mixer, it achieves 58.79% Top-1 (best among compared activations) with negligible computational cost increase.
5. Use TESLA's coefficient budget (A K) as a trainable regularization mechanism
- The improved system can automatically balance expressiveness and generalization by learning appropriate coefficient magnitudes, avoiding overfitting in low-data regimes where KANs and spline methods fail.
6. Replace input-side spectralization (Fourier features) with activation-level spectralization
-
The improved system eliminates the need for handcrafted frequency mappings at the input stage, reducing architectural complexity while achieving better or comparable performance on global interaction tasks like Forrelation.
-
Solve global reasoning problems that require integrating information across many input dimensions simultaneously, such as cryptographic functions, checksum verification, and combinatorial optimization constraints.
-
Learn from highly corrupted data (up to 30% label noise) while maintaining near-Bayes-optimal performance, useful for real-world noisy labeling scenarios.
-
Reconstruct high-fidelity continuous signals (images, PDE solutions, audio) with better accuracy than current state-of-the-art sinusoidal activations, while being more stable during training.
-
Achieve competitive vision performance on standard benchmarks without sacrificing inference speed or parameter efficiency.
-
Adaptively control frequency content during training via learnable coefficients, enabling the network to emphasize high-frequency structure when needed for the task.
-
Scale to larger bit-lengths in discrete problems (up to 27 bits in LPN settings) where other methods fail, suggesting capability for long-range dependency tasks in sequence modeling.
Abstract
The parity problem--deciding whether the number of ones in a binary vector is odd or even--remains challenging for standard neural networks due to linear inseparability and the need for global interactions. We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components. Theoretically, we show that constraining TESLA's coefficients yields Lipschitz/Rademacher complexity bounds and shapes the training dynamics to emphasize higher-frequency structure. Empirically, on parity with input length n = 32, TESLA attains strong generalization with 100K training samples (approximately 0.002% of the 2 32 input space) and remains robust under heavy corruption, retaining high accuracy with up to 30% label noise. We also compare against periodic and frequency-based baselines (SIREN, SNAKE, and Fourier feature embeddings) on parity and Forrelation. Beyond synthetic structure, TESLA delivers comparable performance on ImageNet-100, indicating that activation-level degree control transfers to more general vision workloads. Code: https://github.com/KAU-QuantumAILab/TESLA
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks