Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

arXiv:2608.10288 · cs.LG, cs.CL · Submitted 2026-08-10 · Read on arXiv

Burc Gokden

Fromthesky Research Labs LLC

cs.LG, cs.CL

Submitted: 2026-08-10

Updated: 2026-08-12

Comments: 61 pages, 1 figure, 8 tables

Code: https://github.com/burcgokden/PLDR-LLM-Math-Foundations

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: The paper introduces Power Law Graph Attention (PLGA), a mechanism that generalizes scaled dot-product attention (SDPA) by replacing the fixed bilinear form with a learned, input-generated bilinear

Terminology

Summary

The paper introduces Power Law Graph Attention (PLGA), a mechanism that generalizes scaled dot-product attention (SDPA) by replacing the fixed bilinear form with a learned, input-generated bilinear operator. The architecture, Power Law Decoder Representations (PLDR-LLM), is fully specified and verified against pinned reference releases. Claims are labeled as theorem, conditional theorem, measurement, or conjecture.

Key structural results include: PLGA contains SDPA exactly at the identity operator point; the interaction tensor is strictly entrywise positive with Perron–Frobenius structure; the DAG regularizer has the NOTEARS walk-counting form, with positivity obstructing exact acyclicity; and a commutant criterion identifies operators preserving relative-position dependence under nonresonance.

The paper proves an inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance shows relative fluctuations of 10−6 and below; perturbation bounds quantify but do not certify cached inference. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint.

Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within 5 × 10−5 per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.

The paper's main contributions are: (1) formal definitions of every component and the end-to-end model map, verified against reference implementations; (2) proofs of structural properties underlying three empirical discoveries—operator invariance, a learned singularity condition (rank-one generator), and critical-like training phenomenology; (3) a consolidated comparison of PLDR-LLM with its SDPA base point, with each advantage tied to a result at its actual epistemic strength and costs stated explicitly; (4) collection of open claims as precise, falsifiable conjectures.

The paper also details the architecture: rotary position embeddings with a commutant characterization, the full decoder map, training objective with DAG regularizer, inference with KV-cache and G-cache, and the online generation contract. It discusses power laws as the unique scale-equivariant interaction, criticality of attention dynamics via a conditional spectral dictionary, and the SOC training picture as a phenomenological framework. Advantages over SDPA-LLM include a larger parameterized family with distinct training dynamics, an inspectable law representation, an intrinsic evaluation diagnostic, phase-aware self-instrumented training, inference efficiency, and inductive bias/transfer. Costs include training-time compute and parameters, the criticality search, benchmark parity at small scale, theory debt, and a limited evidence base.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

  1. Replace fixed attention with input-generated bilinear operators – The improved AI can dynamically adapt its attention mechanism per input, capturing higher-order interactions beyond pairwise dot products, enabling better modeling of complex relational data (e.g., multi-hop reasoning, graph-structured knowledge).

  2. Use Perron–Frobenius-positive interaction tensors – The improved AI can guarantee non-negative, stable attention weights with a dominant eigenvalue, preventing vanishing/exploding gradients during training and improving convergence on long sequences.

  3. Apply the commutant criterion for rotary embeddings – The improved AI can preserve relative-position dependencies under nonresonance, allowing more robust positional encoding that maintains performance on sequences longer than those seen in training.

  4. Detect and avoid inference-collapse via invariance monitoring – The improved AI can measure input-invariance of its outputs (relative fluctuations <10−6) and flag when deductive reasoning degenerates to a constant operator, enabling early intervention to maintain genuine inference capability.

  5. Use the DAG regularizer with NOTEARS walk-counting – The improved AI can enforce near-acyclic latent structures during training, improving interpretability and causal inference in generative models, while explicitly acknowledging positivity obstructs exact acyclicity (leading to a controlled relaxation).

  6. Leverage the conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) – The improved AI can implement a measured, stage-wise attention refinement that improves focus on relevant tokens, reducing attention dispersion and improving accuracy on long-context tasks.

  7. Adopt blockwise training under the global Gram – The improved AI can train in blocks while scoring globally, reducing memory overhead without sacrificing answer selection accuracy (verified within 5×10−5 on TruthfulQA), enabling training of larger models on limited hardware.

  8. Use the intrinsic order parameter for self-organized criticality – The improved AI can monitor its own training dynamics (e.g., via spectral properties) and adjust learning rates or architecture in real-time to stay near criticality, improving sample efficiency and robustness to hyperparameter changes.

  9. Implement the learned rank-one singularity condition – The improved AI can detect when its attention operator becomes rank-one (a learned singularity), using this as a diagnostic for mode collapse or over-simplification, and trigger corrective regularization.

  10. Exploit power-law scale-equivariant interactions – The improved AI can generalize across scales (e.g., token frequencies, sequence lengths) by using power-law attention, improving transfer to out-of-distribution lengths and rare-event modeling without retraining.

  11. Use the KV-cache and G-cache with perturbation bounds – The improved AI can perform cached inference with quantified error bounds (not certification), enabling faster decoding while knowing the worst-case deviation, useful for real-time applications with latency constraints.

  12. Apply the theorem-verified components (Lean 4 checked) – The improved AI can include provably correct submodules (e.g., attention operator, regularizer) that are machine-checked, reducing silent bugs in production systems and enabling formal verification of safety-critical AI behaviors.

What the improved AI system can do specifically:

  • Perform multi-hop reasoning on knowledge graphs with dynamic relational attention, outperforming fixed-attention baselines on link prediction and logical entailment.

  • Maintain stable training on sequences 10× longer than training data without positional degradation.

  • Self-diagnose when it is about to collapse into trivial inference (e.g., ignoring input) and automatically re-route to a more expressive mode.

  • Train a 7B-parameter model on a single GPU by blockwise global-Gram scoring, matching full-sequence accuracy within 0.00005 on factual benchmarks.

  • Provide interpretable attention laws (power-law exponents) that humans can inspect to understand why the model made a decision.

  • Operate in real-time with cached inference, knowing the maximum error introduced by caching (e.g., ±0.001 logits).

  • Transfer learned attention patterns from short to long documents without fine-tuning, improving summarization and retrieval on unseen-length inputs.

Abstract

The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator G LM, built from a positive tensor A LM by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Unconditionally: PLGA contains SDPA exactly at G LM=I; A LM and A P are strictly entrywise positive, with Perron-Frobenius structure on A LM; the DAG regularizer has the NOTEARS walk-counting form and positivity obstructs exact acyclicity; and, under nonresonance (satisfied by standard rotary frequencies), a commutant criterion identifies which operators preserve relative-position dependence. An inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance: relative fluctuations of 10-6 and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within 5 times 10-5 per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.

Sources

Related papers