Decodable In-Context State and Model Output Across Training
cs.LG
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- The Geometric Anatomy of Capability Acquisition in Transformers
- Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
- Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express
- Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness
- The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs
- When and How Long? The Readout-Mediator Angle in Temporal Reasoning
- How do Language Models Bind Entities in Context?
- Challenges with unsupervised LLM knowledge discovery
- Inside-Out: Hidden Factual Knowledge in LLMs
- The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior
- Structure-Specific Representational Priors Causally Control the Grokking Delay
- Attention Deficits in Language Models: Causal Explanations for Procedural Hallucinations
- They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It
- Language Models Struggle to Use Representations Learned In-Context
- Eliciting Latent Knowledge from Quirky Language Models
- Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining
- What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- Legible Failures: Detecting and Repairing In-Context Binding Errors
- Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks