Pretraining Transformers with Quantized Softmax in Attention
cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention
- Towards Fully FP8 GEMM LLM Training at Scale
- ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
- ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
- Recipes for Pre-training LLMs with MXFP8
- FP8-LM: Training FP8 Large Language Models
- Theory, Analysis, and Best Practices for Sigmoid Self-Attention
- A Study on ReLU and Softmax in Transformer
- EXAQ: Exponent Aware Quantization For LLMs Acceleration
- Softermax: Hardware/Software Co-Design of an Efficient Softmax for Transformers
- Softmax is not Enough (for Sharp Size Generalisation)
- Replacing softmax with ReLU in Vision Transformers
- BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling
- SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
- SageBwd: A Trainable Low-bit Attention
- IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks