Towards Effective Theory of LLMs: A Representation Learning Approach
cs.LG, cs.AI
Submitted: 2026-05-10
Updated: 2026-09-26
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning
- Revisiting Feature Prediction for Learning Visual Representations from Video
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
- Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
- The Origins of Computational Mechanics: A Brief Intellectual History and Several Clarifications
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Linearity of Relation Decoding in Transformer Language Models
- Prototype-Based Dynamic Steering for Large Language Models
- Priors in Time: Missing Inductive Biases for Language Model Interpretability
- Steering Llama 2 via Contrastive Activation Addition
- Qwen3.5-Omni Technical Report
- Qwen2.5 Technical Report
- Improving Dictionary Learning with Gated Sparse Autoencoders
- Software in the natural world: A computational approach to hierarchical emergence
- Steering Language Models With Activation Engineering
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks