MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries
cs.CL, cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/michelepao1993-dev/attention-experiments
Terminology
Sources
- CoPE: A Lightweight Complex Positional Encoding
- Round and Round We Go! What makes Rotary Positional Encodings useful?
- Longformer: The Long-Document Transformer
- Mistral 7B
- HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models
- Gaussian Error Linear Units (GELUs)
- Training Compute-Optimal Large Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Mixtral of Experts
- Why Language Models Hallucinate
- Scaling Laws for Neural Language Models
- Reformer: The Efficient Transformer
- FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
- Decoupled Weight Decay Regularization
- Large Language Models: A Survey
- Training language models to follow instructions with human feedback
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- RoFormer: Enhanced Transformer with Rotary Position Embedding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering