What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model
Nicolás Vera Zúñiga
Independent Researcher
cs.CL, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: 16 pages, 4 figures. Code, per-run results, and the findings ledger: https://github.com/nicoveraz/token-lattice-ca (archived: https://doi.org/10.5281/zenodo.21880472)
Code: https://github.com/nicoveraz/token-lattice-ca
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper asks what iterated self-feeding probes of language models measure, and answers that they measure "the model and the probe together, in quantities that are not distinguishable by
Terminology
Summary
This paper asks what iterated self-feeding probes of language models measure, and answers that they measure the model and the probe together, in quantities that are not distinguishable by inspection.
The construction is a ring of N token cells, each resampled in place from the model's own windowed conditional pr(xi xi±r) at temperature T, so that the model's output at every site is part of its own input at the next step and no external text enters after initialisation.
The substrate is Glauber dynamics on token sequences, but the change is the coupling: two rings identical except for one flipped token are advanced with common random numbers (CRN) and the same visit order, making undamaged copies diverge by exactly zero, so damage spreading becomes measurable. The token-space Lyapunov exponent λca is the exponential growth rate of damage, and Dnorm is the saturating damage normalised by an independent-noise floor.
The paper's contributions are: (1) a validated instrument, calibrated by reproduction on the Domany–Kinzel automaton where the damage field is provably the automaton itself, matching an independent prediction bit-exactly with zero mismatching cells; (2) a manufactured phase transition that belongs to the probe, measurable to three decimal places, with a mechanism (an attracting fixed point of the argmax map), a boundary (radius r ∈ 1,2 only), and a control (a masked-LM construction with no such fixed point shows no transition); (3) a discriminator that separates construction-determined from model-determined readings by varying one factor with the other held fixed; (4) estimator gating, reporting four retracted verdicts each caught by a known-answer system.
The discriminator is tabulated with five manipulations. Varying the construction moves the instrument (the transition disappears); varying the model across 19 models and 70× scale does not move it; varying family and construction together separates the readouts — λca does not survive it and the attractor share does; varying the training checkpoint moves λca (it crosses zero at a reproducible point in training); ablating internal components moves ignition by up to 0.33.
Construction-determined readings include the damage light cone, which is kinematic (its extent is set by the update window rather than the model), and the radius scaling λca(r), which is model-invariant across 19 models and two scale ladders spanning 70×. Model-determined readings include λca crossing zero at a reproducible point in training (e.g., −0.0185 at step 256, +0.0679 at step 512, +0.1923 at step 1000, plateauing around +0.1558 to +0.1792 at later steps), and the attractor share, which ranks models consistently across constructions (ρ = +0.752) and is seed-stable at 0.848.
The paper reports a phase transition measured to three decimal places that belongs to the probe rather than to any language model, and notes the authors themselves mistook it for a model property for four months. The test that separates them is: hold the construction fixed and vary the model, or hold the model fixed and vary the construction, and see which readings move.
The estimator gating section reports four retracted verdicts, each caught by a known-answer system rather than by review: a calibration run at a geometry the measurement never used; correlated replicas from a shared visit order shrinking error bars 8×; an unbounded minimum landing on the edge of a scan; and a test never shown able to discriminate. A further defect class is identified: a statistically-shaped criterion applied to a quantity with no room to vary, including a correlation function that ranked a constant vector (all values exactly 0.000) by input position, producing ρ = +0.829, p = 0.058.
Limits include: the developmental transition is single-family (Pythia); cross-family, loss does not organise it better than tokens; the explanandum is not internal — λca is largely fixed by the state the model drives the lattice into, with four routes failing to attach it to a named internal mechanism; the dynamics are not reducible to annealed mean field; and scope is bounded to greedy decoding, one radius for ablation work, and one architecture family for component manipulations.
The paper concludes that iterated self-feeding probes mix construction-determined and model-determined quantities in readings that look alike, and the test in §5 separates them. It resists the summary that the instrument measures itself, stating instead that "the instrument reads the model. It also reads the probe, in quantities that carry the same units and the same apparent precision, and the cost of not checking which is which is four months and a retracted phase transition." The sharper statement is that which of its quantities reads the model is itself something to be measured rather than assumed, and the answer is not uniform across them.
Improvements for AI systems
Improvements to AI systems:
-
Construction-aware self-feedback calibration: Add a discriminator module to any iterated self-feeding or self-distillation loop (e.g., recursive prompting, chain-of-thought self-consistency, or self-training) that systematically varies the construction (e.g., prompt template, sampling window, update order) while holding the model fixed, and vice versa. The improved system can flag readings that move with construction as probe artifacts, preventing false conclusions about model capabilities.
-
Known-answer gating for retracted verdicts: Implement an automated estimator gate that runs every reported metric against a set of known-answer baselines (e.g., constant inputs, random noise, degenerate geometries) before accepting it. The improved system can reject statistically-shaped criteria applied to quantities with no variance, catching false positives like the ρ = +0.829 correlation on a constant vector.
-
Common-random-number damage tracking for stability audits: Use paired replicas with identical random seeds and visit orders (CRN) to measure divergence in any iterative AI system (e.g., multi-agent simulations, evolutionary algorithms, or recurrent generation). The improved system can distinguish kinematic/construction-determined divergence from model-determined divergence, and report a Lyapunov-style exponent to quantify sensitivity to single-token perturbations.
-
Phase-transition attribution testing: Before claiming a phase transition or emergent behavior in a language model, run the two-factor test: hold construction fixed and vary model (across families and scales), then hold model fixed and vary construction. The improved system can classify any observed transition as either model-determined (survives model variation) or probe-determined (moves with construction), avoiding the four-month error the authors made.
-
Checkpoint-sensitive early-training diagnostics: Use the λca crossing-zero metric (e.g., −0.0185 at step 256, +0.0679 at step 512) as a reproducible training-progress indicator. The improved system can monitor a token-space damage exponent during training to detect when a model begins to organize its output state, providing a construction-independent signal for early stopping or curriculum adjustments.
-
Attractor-share ranking with cross-construction stability: Compute the attractor share (fraction of sites converging to a fixed point) across multiple constructions and use its rank correlation (ρ = +0.752) as a robust model-comparison metric. The improved system can rank models consistently even when other metrics (like λca) are construction-sensitive, enabling fairer model evaluation.
-
Masked-LM control for fixed-point artifacts: When a probe shows an attracting fixed point (e.g., argmax collapse), run a control with a masked-LM construction lacking that fixed point. The improved system can identify whether a measured behavior is an artifact of the probe's update rule rather than the model's intrinsic property, and adjust the probe design accordingly.
-
Seed-stability verification for all reported numbers: For any AI system output, report seed-stability (e.g., 0.848) alongside point estimates. The improved system can flag metrics that are not seed-stable as unreliable, and automatically retract or re-estimate them with variance-aware reporting.
-
Kinematic light-cone correction: When measuring information propagation or influence in a recurrent or self-feeding system, subtract the kinematic contribution set by the update window (radius r) from the measured damage spread. The improved system can isolate model-driven dynamics from trivial window-limited propagation, yielding cleaner causal attributions.
-
Cross-family loss-token organization check: Before using loss as a proxy for internal organization, test whether it organizes better than token count across families. The improved system can avoid misattributing developmental trends to loss when token count is the actual driver, as shown by the single-family limitation.
Abstract
A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction chosen to make the question sharp: a ring of token cells resampled in place by the model's own windowed conditional p r(x i x i+-r). The substrate is Glauber dynamics on token sequences and is not new; what we change is the coupling. Advancing two rings that differ in one token under common random numbers makes undamaged copies diverge by exactly zero, so damage spreading becomes measurable where a maximal coupling gives mixing times instead. The answer is that it measures two different things at once, in readings that look alike. Some quantities are fixed by the construction: the damage light cone is kinematic, and the radius scaling of the token-space Lyapunov exponent lambda ca(r) is model-invariant across 19 models and two scale ladders spanning 70x. Others genuinely track the model: lambda ca crosses zero at a reproducible point in training, and the attractor share ranks models consistently however the lattice is built. Left undistinguished, the first kind is readily mistaken for the second -- we did so ourselves for four months, and report a phase transition we measured to three decimal places that belongs to the probe rather than to any language model. We give the test that separates them: hold the construction fixed and vary the model, or hold the model fixed and vary the construction, and see which readings move. We validate the instrument by reproduction first, recovering a Domany-Kinzel damage field bit-exactly against an independent prediction, and we report the estimator failures that this discipline caught -- four retracted verdicts, each on a quantity that looked like a measurement. The methodology ships as a package.
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering