FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward
cs.LG, cs.DC
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/srimanachanta/fold-attention
Terminology
Sources
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- The Llama 3 Herd of Models
- Hydragen: High-Throughput LLM Inference with Shared Prefixes
- Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding
- FP8 Formats for Deep Learning
- Online normalizer calculation for softmax
- gpt-oss-120b & gpt-oss-20b Model Card
- DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
- SparQ Attention: Bandwidth-Efficient LLM Inference
- Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
- VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
- Qwen3 Technical Report
- FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
- BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
- Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
- Diagnosing Training Inference Mismatch in LLM Reinforcement Learning via a Zero-Mismatch Reference
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks