The Gradient Does Not See Rank: Rank-Indifference in Matrix-CODI on ProsQA
cs.LG
Submitted: 2026-09-02
Updated: 2026-09-02
Terminology
Sources
- Implicit Regularization in Deep Learning May Not Be Explainable by Norms
- The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
- Implicit Regularization in Deep Matrix Factorization
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
- LLM Latent Reasoning as Chain of Superposition
- SIM-CoT: Supervised Implicit Chain-of-Thought
- Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
- State Rank Dynamics in Linear Attention LLMs
- Continuous Chain of Thought Enables Parallel Exploration and Reasoning
- Addressing Token Uniformity in Transformers via Singular Value Transformation
- Implicit Regularization in Matrix Factorization
- Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
- Training Large Language Models to Reason in a Continuous Latent Space
- Weight decay induces low-rank attention layers
- Optimal ablation for interpretability
- Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure
- Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
- The Key to State Reduction in Linear Attention: A Rank-based Perspective
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks