In-Context Binding Capacity in Language Models
cs.LG
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
- Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
- Birth of a Transformer: A Memory Viewpoint
- How do Language Models Bind Entities in Context?
- When Attention Sink Emerges in Language Models: An Empirical View
- Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
- Working Memory Capacity of ChatGPT: An Empirical Study
- Training Compute-Optimal Large Language Models
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Scaling Laws for Neural Language Models
- Code Pretraining Improves Entity Tracking Abilities of Language Models
- Comparing Transformers and Hybrid Models at the Token Level
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- Superposition Yields Robust Neural Scaling
- Lost in the Middle: How Language Models Use Long Contexts
- Scaling Data-Constrained Language Models
- In-context Learning and Induction Heads
- Are Emergent Abilities of Large Language Models a Mirage?
- Massive Activations in Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks