Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers
cs.LG
Submitted: 2026-08-31
Updated: 2026-08-31
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Towards Combinatorial Interpretability of Neural Computation
- On the Bottleneck of Graph Neural Networks and its Practical Implications
- Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
- On the Design Space Between Transformers and Recursive Neural Nets
- The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
- A Generalization of Transformer Networks to Graphs
- Looped Transformers for Length Generalization
- Looped Transformers as Programmable Computers
- Contextual Position Encoding: Learning to Count What's Important
- On Oversquashing in Graph Neural Networks Through the Lens of Dynamical Systems
- Universal Length Generalization with Turing Programs
- A Formal Framework for Understanding Length Generalization in Transformers
- Length Generalization in Arithmetic Transformers
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Functional Interpolation for Relative Positions Improves Long Context Transformers
- What graph neural networks cannot learn: depth vs width
- Transformers Can Do Arithmetic with the Right Embeddings
- A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks