More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations
cs.LG, cs.AI, stat.ML
Submitted: 2026-05-26
Updated: 2026-08-28
Code: https://github.com/karpathy/nanoGPT
Terminology
Sources
- Learning Activation Functions to Improve Deep Neural Networks
- Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
- Learning Activation Functions: A new paradigm for understanding Neural Networks
- Gaussian Error Linear Units (GELUs)
- Training Compute-Optimal Large Language Models
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- Adam: A Method for Stochastic Optimization
- Muon is Scalable for LLM Training
- KAN: Kolmogorov-Arnold Networks
- Decoupled Weight Decay Regularization
- Learning Combinations of Activation Functions
- Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
- Mish: A Self Regularized Non-Monotonic Activation Function
- GLU Variants Improve Transformer
- Adaptive Blending Units: Trainable Activation Functions for Deep Neural Networks
- LLaMA: Open and Efficient Foundation Language Models
- Quick and (not so) Dirty: Unsupervised Selection of Justification Sentences for Multi-hop Question Answering
- Qwen2 Technical Report
- HellaSwag: Can a Machine Really Finish Your Sentence?
- ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks