Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks
cs.LG
Submitted: 2026-07-29
Updated: 2026-09-28
Comments: 18 pages, 5 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Since GPT, most Transformers have repeated the same attention mechanism at every layer.
Terminology
Abstract
Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choice, placement, or both, making causal attribution difficult. We introduce Aether-7B-5Attn, a 6.59B-parameter mixture-of-experts model (about 2.98B active) whose 49 layers contain seven sequence-mixing mechanisms arranged as a 7 times7 Latin square. Because each mechanism appears exactly once in every row and column, the design guarantees balanced exposure across depth while eliminating placement confounds. To evaluate this principle, we build a parameter-matched proxy with four mechanisms arranged as a 4 times4 Latin square over sixteen layers, matched to 700.9M parameters and trained with eight seeds per arm. The results reveal a clear dissociation. Rearranging a distributed heterogeneous stack into a balanced periodic cycle changes validation loss by only 0.16%, indicating that exact placement has little effect. In contrast, clustering the same mechanisms into contiguous depth bands incurs a 0.59% penalty, while replacing the heterogeneous stack with a homogeneous one incurs a 1.68% penalty. These results indicate that performance depends primarily on heterogeneous composition distributed across depth rather than on any particular permutation. We confirm this finding at 2.16 times larger scale (1.514B parameters), where the homogeneous-stack penalty increases to 2.63% and removing the SSM-family mechanism produces a 3.20% degradation. We further report per-mechanism cost profiles, English and Korean evaluations, and a causal-safety audit of all 49 layers. We release model weights, training recipes, training code, logs, and architecture source code.
Sources
- PaLM: Scaling Language Modeling with Pathways
- Training Compute-Optimal Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Longformer: The Long-Document Transformer
- Linformer: Self-Attention with Linear Complexity
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Retentive Network: A Successor to Transformer for Large Language Models
- RWKV: Reinventing RNNs for the Transformer Era
- xLSTM: Extended Long Short-Term Memory
- Jamba: A Hybrid Transformer-Mamba Language Model
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks