Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
Burc Gokden
Fromthesky Research Labs LLC
cs.LG, cs.CL
Submitted: 2026-08-10
Updated: 2026-08-12
Comments: 61 pages, 1 figure, 8 tables
Code: https://github.com/burcgokden/PLDR-LLM-Math-Foundations
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: The paper introduces Power Law Graph Attention (PLGA), a mechanism that generalizes scaled dot-product attention (SDPA) by replacing the fixed bilinear form with a learned, input-generated bilinear
Terminology
Summary
The paper introduces Power Law Graph Attention (PLGA), a mechanism that generalizes scaled dot-product attention (SDPA) by replacing the fixed bilinear form with a learned, input-generated bilinear operator. The architecture, Power Law Decoder Representations (PLDR-LLM), is fully specified and verified against pinned reference releases. Claims are labeled as theorem, conditional theorem, measurement, or conjecture.
Key structural results include: PLGA contains SDPA exactly at the identity operator point; the interaction tensor is strictly entrywise positive with Perron–Frobenius structure; the DAG regularizer has the NOTEARS walk-counting form, with positivity obstructing exact acyclicity; and a commutant criterion identifies operators preserving relative-position dependence under nonresonance.
The paper proves an inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance shows relative fluctuations of 10−6 and below; perturbation bounds quantify but do not certify cached inference. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint.
Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within 5 × 10−5 per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.
The paper's main contributions are: (1) formal definitions of every component and the end-to-end model map, verified against reference implementations; (2) proofs of structural properties underlying three empirical discoveries—operator invariance, a learned singularity condition (rank-one generator), and critical-like training phenomenology; (3) a consolidated comparison of PLDR-LLM with its SDPA base point, with each advantage tied to a result at its actual epistemic strength and costs stated explicitly; (4) collection of open claims as precise, falsifiable conjectures.
The paper also details the architecture: rotary position embeddings with a commutant characterization, the full decoder map, training objective with DAG regularizer, inference with KV-cache and G-cache, and the online generation contract. It discusses power laws as the unique scale-equivariant interaction, criticality of attention dynamics via a conditional spectral dictionary, and the SOC training picture as a phenomenological framework. Advantages over SDPA-LLM include a larger parameterized family with distinct training dynamics, an inspectable law representation, an intrinsic evaluation diagnostic, phase-aware self-instrumented training, inference efficiency, and inductive bias/transfer. Costs include training-time compute and parameters, the criticality search, benchmark parity at small scale, theory debt, and a limited evidence base.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
-
Replace fixed attention with input-generated bilinear operators – The improved AI can dynamically adapt its attention mechanism per input, capturing higher-order interactions beyond pairwise dot products, enabling better modeling of complex relational data (e.g., multi-hop reasoning, graph-structured knowledge).
-
Use Perron–Frobenius-positive interaction tensors – The improved AI can guarantee non-negative, stable attention weights with a dominant eigenvalue, preventing vanishing/exploding gradients during training and improving convergence on long sequences.
-
Apply the commutant criterion for rotary embeddings – The improved AI can preserve relative-position dependencies under nonresonance, allowing more robust positional encoding that maintains performance on sequences longer than those seen in training.
-
Detect and avoid inference-collapse via invariance monitoring – The improved AI can measure input-invariance of its outputs (relative fluctuations <10−6) and flag when deductive reasoning degenerates to a constant operator, enabling early intervention to maintain genuine inference capability.
-
Use the DAG regularizer with NOTEARS walk-counting – The improved AI can enforce near-acyclic latent structures during training, improving interpretability and causal inference in generative models, while explicitly acknowledging positivity obstructs exact acyclicity (leading to a controlled relaxation).
-
Leverage the conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) – The improved AI can implement a measured, stage-wise attention refinement that improves focus on relevant tokens, reducing attention dispersion and improving accuracy on long-context tasks.
-
Adopt blockwise training under the global Gram – The improved AI can train in blocks while scoring globally, reducing memory overhead without sacrificing answer selection accuracy (verified within 5×10−5 on TruthfulQA), enabling training of larger models on limited hardware.
-
Use the intrinsic order parameter for self-organized criticality – The improved AI can monitor its own training dynamics (e.g., via spectral properties) and adjust learning rates or architecture in real-time to stay near criticality, improving sample efficiency and robustness to hyperparameter changes.
-
Implement the learned rank-one singularity condition – The improved AI can detect when its attention operator becomes rank-one (a learned singularity), using this as a diagnostic for mode collapse or over-simplification, and trigger corrective regularization.
-
Exploit power-law scale-equivariant interactions – The improved AI can generalize across scales (e.g., token frequencies, sequence lengths) by using power-law attention, improving transfer to out-of-distribution lengths and rare-event modeling without retraining.
-
Use the KV-cache and G-cache with perturbation bounds – The improved AI can perform cached inference with quantified error bounds (not certification), enabling faster decoding while knowing the worst-case deviation, useful for real-time applications with latency constraints.
-
Apply the theorem-verified components (Lean 4 checked) – The improved AI can include provably correct submodules (e.g., attention operator, regularizer) that are machine-checked, reducing silent bugs in production systems and enabling formal verification of safety-critical AI behaviors.
What the improved AI system can do specifically:
-
Perform multi-hop reasoning on knowledge graphs with dynamic relational attention, outperforming fixed-attention baselines on link prediction and logical entailment.
-
Maintain stable training on sequences 10× longer than training data without positional degradation.
-
Self-diagnose when it is about to collapse into trivial inference (e.g., ignoring input) and automatically re-route to a more expressive mode.
-
Train a 7B-parameter model on a single GPU by blockwise global-Gram scoring, matching full-sequence accuracy within 0.00005 on factual benchmarks.
-
Provide interpretable attention laws (power-law exponents) that humans can inspect to understand why the model made a decision.
-
Operate in real-time with cached inference, knowing the maximum error introduced by caching (e.g., ±0.001 logits).
-
Transfer learned attention patterns from short to long documents without fine-tuning, improving summarization and retrieval on unseen-length inputs.
Abstract
The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator G LM, built from a positive tensor A LM by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Unconditionally: PLGA contains SDPA exactly at G LM=I; A LM and A P are strictly entrywise positive, with Perron-Frobenius structure on A LM; the DAG regularizer has the NOTEARS walk-counting form and positivity obstructs exact acyclicity; and, under nonresonance (satisfied by standard rotary frequencies), a commutant criterion identifies which operators preserve relative-position dependence. An inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance: relative fluctuations of 10-6 and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within 5 times 10-5 per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.
Sources
- Layer Normalization
- XCiT: Cross-Covariance Image Transformers
- On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning
- CoulGAT: An Experiment on Interpretability of Graph Attention Networks
- Power Law Graph Transformer for Machine Translation and Representation Learning
- PLDR-LLM: Large Language Model from Power Law Decoder Representations
- PLDR-LLMs Learn A Generalizable Tensor Operator That Can Replace Its Own Deep Neural Net At Inference
- PLDR-LLMs Reason At Self-Organized Criticality
- HyperNetworks
- Going Beyond Linear Transformers with Recurrent Fast Weight Programmers
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Bilinear Attention Networks
- Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Pointer Sentinel Mixture Models
- LieRE: Lie Rotational Positional Encodings
- DBA: Efficient Transformer with Dynamic Bilinear Low-Rank Attention
- Sinkformers: Transformers with Doubly Stochastic Attention
- Linear Transformers Are Secretly Fast Weight Programmers
- Attention as a Hypernetwork
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks