Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures
cs.CL, cs.AI
Submitted: 2026-09-06
Updated: 2026-09-06
Comments: 72 pages, 13 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: "An agent that writes its conclusions into a store it later retrieves from closes a loop usually reported as one-way contamination.
Terminology
Abstract
"An agent that writes its conclusions into a store it later retrieves from closes a loop usually reported as one-way contamination. Taking the loop to the infinite-tenure limit against an append-only store gives a different picture: because writing never deletes, the reachable state space has a hard upper edge at (n-1)/n, so the outcome is a choice between two edges rather than a decay. At f 0 = 0.9 the interval between the two modes holds 3.6% of 220 runs where a uniform spread would put 20.6%, and is strictly empty on the first 15; the pooled mean describes 8.2% of the runs it summarises, the median 68.2%. Everything the model contributes is carried by one measured primitive with no fitted parameter, the copy function γ(ϕ): on 36 Wikidata facts, sign(γ - γ crit), with γ crit = 1/k at r = 0, w = 1, predicts the direction of drift on 353 of 360 real-fact runs (39 of 40 synthetic in the same batch). Scale does not rescue the store: pooled frontier capture is 0.850, with claude-sonnet-4.5 captured on 20 of 20 seeds against our registered prediction of <0.5. What the interval tests is distinguishability rather than count: on the real facts, multi-valued runs have 6.4x its occupancy of the rest. It survives at f 0 in 0.1, 0.3, 0.5, capture peaks at f 0 = 0.5, and of four interventions with criteria frozen first, timing dominates fraction at matched budget while a consistency gate drives every model to 0.993. The resampling unit is the seed, at a design effect of 3.75 on a pooled level: under a 44-seed control the ordering supporting claim 4 collapses from Spearman +0.98 at three seeds to +0.31-0.80 at forty-four, while claim 2's ordering is exact there (+1.00, p = 0.017). All 87 graded rows are in Appendix W, 37 of them graded withdrawn, failed, self-correcting, undecidable or an acknowledged limit, against 50 that are not."
Sources
- Self-Consuming Generative Models Go MAD
- RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Chinchilla Scaling: A replication attempt
- Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop
- STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
- The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- Universality of the $\pi^2/6$ Pathway in Avoiding Model Collapse
- Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents
- MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents
- Memory Injection Attacks on LLM Agents via Query-Only Interaction
- Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
- Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval
- A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
- Self-Correcting Self-Consuming Loops for Generative Model Training
- The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text
- MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models
- Minority Sentinel: When to Overturn Majority Voting in Multi-Agent LLM Debates
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering