Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction
cs.LG, cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/burcgokden/PLDR-LLM-Training-Dynamicshttps:
Terminology
Sources
- Layer Normalization
- On the Expressivity Role of LayerNorm in Transformers' Attention
- Adaptive Gradient Methods at the Edge of Stability
- Setting the Record Straight on Transformer Oversmoothing
- CoulGAT: An Experiment on Interpretability of Graph Attention Networks
- PLDR-LLMs Learn A Generalizable Tensor Operator That Can Replace Its Own Deep Neural Net At Inference
- PLDR-LLMs Reason At Self-Organized Criticality
- Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
- LoRA: Low-Rank Adaptation of Large Language Models
- An exact mapping between the Variational Renormalization Group and Deep Learning
- The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
- On the Role of Attention Masks and LayerNorm in Transformers
- Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit
- Stabilizing Transformer Training by Preventing Attention Entropy Collapse
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks