Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding
Kushal Chakrabarti
South Park Commons
cs.AI, cs.LG, cs.SE
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale.
Terminology
Summary
Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression costs O(2D) in a prompt of D instructions. We name the resulting divergence catastrophic remembering, the inverse of catastrophic forgetting around which continual learning is organized. First, we characterize this phenomenon across 247,694 instruction lifetimes in 1,867 repositories: agentic prompts grow without bound, more than tripling over their lifetime (+226%), gaining +4.9 net instructions every commit; further, the older an instruction gets, the less likely it is to be deleted (log-hazard-0.032/commit). Then, we show that prompt comments can halt the growth: inverting IFEval yields verifiable worlds whose optimal prompts are known, and there comments encoding latent reasoning remove 99.3% of excess instructions (+211.3% to +1.4%). Finally, applying the same inversion to WildIFEval, we show that prompt comments can improve real-world agentic instruction-following by up to 23.1%. If English is the new code, why don't we have comments yet?
Overwhelmingly, agentic READMEs never stop growing. A copilot-instructions.md, AGENTS.md or CLAUDE.md gains instructions and rarely shed them. Deleting the file is not the answer: repositories carrying one finish agent tasks faster and in fewer tokens. Keeping everything is not free either: instruction-following degrades as constraints accumulate, and the median file already carries 39 instructions, the average more than tripling over its own lifetime (+226%).
Different mechanisms can explain this growth. If requirements stop applying (instruction staleness), deletion hazard rises with an instruction's age. If brittle instructions die young, leaving robust ones behind (content fragility), hazard falls. If maintainers lose an instruction's rationale (imperfect recall), hazard falls with age and, unlike either rival, also with the number of maintainers. Nobody has identified the root cause because line diffs and file sizes cannot resolve individual instructions over time.
What's the point of this code?
is a question every engineer has asked. Deleting agentic instructions risks a regression, and doing so safely requires counterfactuals that maintainers cannot feasibly run. Recording an instruction's latent reasoning costs O(1), but reconstructing it costs O(2D) for a prompt with D instructions. Successive edits by multiple authors destroy that latent reasoning, so deletion hazard decays with age, additions outrun removals, and instructions grow without bound until agents become non-functional — catastrophic remembering. Under catastrophic forgetting, a gradient learner overwrites what it should have kept; under catastrophic remembering, a maintainer keeps what they should have overwritten. Both fail and fail for the same reason: whatever would license the update is gone.
How deletions actually occur in agentic prompts points to the root cause. 76.8% of instruction deaths arrive in one commit that bulldozes the file, after which growth resumes at its prior rate: removing one instruction needs a reason, removing all needs none. And the hazard agrees — removal becomes less likely the older an instruction is, and less likely still the more authors have edited the file.
Software engineering solved this decades ago with the comment: text addressed to the next maintainer and invisible to the interpreter. Recovering why code was written as it was is developers' most serious problem, so the practice is to record the why, not the how or the what. Unsurprisingly, both carry over to agentic prompts.
We show that prompt comments encoding an instruction's latent reasoning can halt unbounded prompt growth, removing 99.3% of the excess size (+211.3% to +1.4%) over 51 steps, and cutting it from +60.4% to −5.8% (66.2pp) over 15. Because irrelevant context distracts models, this improves instruction-following on real prompts by as much as 23.1% (11.6pp). We ablate the effect across comment payloads to show outcome-grounded latent reasoning drives the gain. Finally, we show that prompt comments' advantage only grows as (agentic) maintainers grow more capable.
We contribute a new approach to prompt maintenance, characterized across four dimensions:
• Theory of prompt maintenance. Decaying recoverability of latent reasoning predicts unbounded growth.
• Root cause of unbounded prompt growth. Deletion hazard decays with age over 247,694 lifetimes, implicating imperfect recall, not instruction staleness or content fragility.
• New evaluation with known optimum. Inverting IFEval yields verifiable worlds whose minimum cover is known, so excess size and correctness become measurable.
• Practical solution. Prompt comments remove 99.3% of the excess size, and improve instruction-following on real prompts by as much as 23.1%.
We model prompt maintenance as the online estimation of an unobservable constraint set from censored, noisy feedback, and derive an equilibrium prompt size that diverges as write-time provenance decays. Without that provenance, the optimal estimator becomes append-only.
A task j = (oj, Cj) pairs a stated objective oj with its unobservable constraints Cj ⊆ C, which are drawn from a fixed hidden set C = gc of verifiers. The prompt at time t is a finite set of instructions Dt, with instruction d entering the prompt at time τd and having age ad = t − τd. A response to j satisfies c ∈ Cj with probability qc(D, j); C is unobservable except for draws from qC.
At each timestep t, three parties interact in the system. A (possibly agentic) maintainer adds and deletes instructions, updating the prompt Dt. The harness draws tasks j, runs them using the executor under prompt Dt, and then verifies them using Cj, generating draws from qC.
A maintainer's target is the minimum cover D⋆, the smallest instruction set that maximizes expected constraint satisfaction. Formally, D⋆ = argmin D, D∈M, M = argmax Ej, c∈Cj [qc(D, j)]. Every experiment in this paper scores Dt against D⋆ on two axes: correctness, the expected satisfaction E[qc], and excess size, Dt/D⋆ − 1. A maintenance regime wins only by improving one axis without degrading the other.
Without assuming a particular functional form, we assume qC meets the properties listed in Table 1. Properties A1–A3 (instructability, interference, and redundancy) set the overall regime whereas properties A4 (censoring) and A5 (stochasticity) respectively block deletions and drive additions.
Deleting instructions without risking a regression requires a maintainer to run an infeasible number of counterfactuals. Write d's contribution to c as ∆c(d D) = Ej[qc(D, j) − qc(D d, j)] over tasks with c ∈ Cj; d is excess when ∆c ≤ 0 for every c. Estimating it means probing D d, and redundancy defeats one-at-a-time probes: two instructions covering one constraint each look free alone, though deleting both breaks it. An honest audit therefore costs O(2Dt) subset probes, and D⋆ is incomputable in general. Knowing why a maintainer added d — its latent reasoning rd — collapses that to O(1), but rd rapidly decays with ad.
Recoverability thus sets equilibrium size. Define the decay ρ(d, a) = Pr(rd recoverable at age a), with marginal ρ̄(a) = Ed ρ(d, a); thus, ρ(d, 0) = ρ̄(0) = 1 and both ρ(d, a) → 0 and ρ̄(a) → 0 as a → ∞. We call that decay imperfect recall: the instruction stays and its failures recur, but its latent reasoning is increasingly unrecoverable. A maintainer rationally deletes an instruction only when that reason survives and it verifies as excess, so the hazard in its own age factorizes (the event-history form used for written organizational rules), h(a) ≈ ρ̄(a) s(a), for s(a) = Pr(excess a). Additions net against the summed hazard, E[Dt+1 − Dt] = At − Σd h(ad), for episode additions At. Flow balance then yields unbounded growth: D∞ = A / (ρ̄(a) s) → ∞ as a → ∞.
Equation 5 is catastrophic remembering in closed form: excess size diverges as ρ̄ → 0, even with additions arriving at a constant rate and fixed constraints. That is, even if the task does not change, the decaying memory of why each instruction was written is sufficient to drive the unbounded growth of instructions.
Recoverability is therefore a necessary intervention and, as we show later, also a sufficient one. If you let latent reasoning decay, the only practical deletion left available is the wholesale rewrite; if you restore it, the prompt settles near D⋆.
Agentic context files grow until someone rewrites them wholesale and then they grow again, characterizing the unique sawtooth-like growth curve we call the ratchet. Its deletion hazard and covariates identify its root cause for the first time — imperfect recall.
We decompose agentic prompts into individual instructions and track their decay over commit histories, then falsify instruction staleness and content fragility and finally confirm a prediction only imperfect recall can make. Our corpus spans 1,867 GitHub repositories, 1,801 multi-version files, 299,440 version-to-version transitions, and 247,694 instruction lifetimes.
Maintainers add and almost never remove, so agentic prompts grow without bound in instruction count, instruction complexity and total size until history ends or a file is rewritten wholesale. At its last tracked version, the median file holds 39 instructions (90th percentile: 131), well past the threshold at which instruction-following degrades. Among 1,576 multi-version repositories, 64.3% grow their instruction count against 26.6% that shrink it, a median net gain of +7 instructions. On average, excluding mass rewrites, each commit adds a net +4.9 instructions across 19,267 commits.
Every component we measure grows. Over a file's own lifetime the mean trajectory gains up to +226% in count and +10% in mean instruction length. Alongside that, the total size of agentic prompts grows +140%, so the count is not text migrating between the instruction and payload classes.
When maintainers delete instructions, they generally do so wholesale in a mass rewrite. 77.3% of instruction deaths arrive in a wholesale rewrite or a migration to a sibling file. In such cases, a line diff cannot separate deletions from rewordings, which partially explains why deletions are rare
has stood unexplained: prompt-change studies report additions dominating without settling what became of any one instruction. When a file loses half or more of its instructions in a single commit, we call it a rewrite; text reappearing in a sibling file is a migration. We censor both as competing risks, since neither involves a maintainer judging any given instruction to be worth its place. Matching validates at 1.000 precision and 0.933 recall on 50 hand-annotated transitions, leaving 247,694 lifetimes and 28,426 tracked deletions.
Although a mass rewrite resets a prompt's size, it does not address the underlying root cause (imperfect recall), so the prompt starts growing again immediately after the rewrite. Aligning files on their first mass rewrite at t = 0, the mean instruction count rises monotonically through t = −1, drops to 59.5% of its pre-rewrite value at t = 0, and recovers to 91.5% over the 10 commits after. In fact, agentic prompts grow faster after a rewrite: a file gains 4.9% instructions per commit after a mass rewrite vs. 4.1% before it.
The deletion hazard h(a) and its interaction with author count identify imperfect recall as the mechanism driving instruction growth. In particular, instruction staleness and imperfect recall predict opposite slopes for deletion hazard: the first predicts hazard rising with age, as instructions fall out of date; the second predicts it falling, as the reasoning behind them is lost. As we can see in Figure 5, the deletion hazard h(a) falls with age a. The log-hazard slope from a repository-stratified bootstrap is −0.032 per commit, its interval excluding zero (95% CI [−0.047, −0.019]).
Content fragility — where fragile instructions, files or repositories die young — would bend deletion hazard downward through compositional effects, but composition alone cannot account for the slope. We refit it with a gamma frailty model (n = 28,255 deaths, censored at 50 commits) against a no-frailty baseline. Its strongest form, shared instructions with the same text, absorbs 30.8% but leaves the slope at −0.0355 ([−0.0414, −0.0296]).
Finally, imperfect recall predicts something that neither rival can explain: if latent reasoning is encoded while a maintainer makes the change, deletion hazard should decay with the number of maintainers touching a file, not just with age. Counting human authors and controlling for file activity, it does: βmulti-author×age = −0.021 (z = −11.7).
Across real-world repositories, then, unbounded prompt growth is widespread, costly at the sizes files reach, and explainable only by imperfect recall. However, this setting is observed, not assigned; we turn next to one where we control what the maintainer inherits.
Prompt comments can permanently halt the unbounded growth of instructions in agentic prompts and, in doing so, improve instruction-following through smaller, more robust prompts. To evaluate prompt effectiveness, we first propose a new evaluation for agentic prompts in task-stationary settings. We use this evaluation to demonstrate that prompt comments enable prompts to settle near their theoretically optimal minimum cover. Finally, we show that prompt comments can buy back instruction-following lost to extraneous, noisy instructions in real-world prompts.
In practice, maintainers in agentic coding contexts must learn an effective prompt to satisfy a set of unobservable constraints. In order to maximize correctness, this prompt must simultaneously cover all constraints while, given well-known issues with instruction-following decay, being minimal. Evaluating a prompt therefore requires measuring both its excess size and its correctness, but calculating excess size requires minimum cover, which is incomputable in general. Standard instruction-following suites offer a path.
By inverting IFEval, we can attack the problem. For each benchmark item (Dj, Cj) — instructions Dj, verifiers Cj — we apply the following transform:
-
Hide Dj. The item's stated instructions become the reference minimum cover D⋆, so D⋆ is known a priori and excess size is measurable.
-
Keep Cj. Its verifiers become the world's hidden constraint set, run by the harness alone and never named to the maintainer.
-
Generate brief oj. A stronger model lossily summarizes Dj as a general task objective (
draft a product announcement
).
A fresh maintainer then reconstructs Dj from oj over T steps, seeing only censored, noisy feedback from Cj. Its arm decides whether it may annotate an instruction d with a comment rd carrying that instruction's latent reasoning. The next maintainer reads Dt and every rd; the executor reads only Dt. An instruction tells the executor what to do; a comment tells the next maintainer why. Enforcing that split, the harness strips comments before the prompt reaches the executor and rejects any instruction citing a past failure, leaving the comment as the only surviving channel for latent reasoning.
Prompt comments encoding an instruction's latent reasoning settle the prompt at its minimum cover. Its latent reasoning summarizes the failure behind its instruction, a hypothesis, and how it has fared. Concretely, across 552 maintenance histories, maintainers encoding latent reasoning as prompt comments learn prompts at −5.8% excess size against +60.4% for maintainers without them (66.2pp), at parity constraint satisfaction. Arms differ only in the handoff: the prompt instructions along with its comments (or not), containing a maintainer's summarized latent reasoning.
As (agentic) maintainers scale, they ratchet harder and prompt comments help them more. Varying the maintainer across 3 tiers at T=15, the uncommented arm's excess rises from +67.7% to +571.9%; at the top tier, prompt comments enable strictly Pareto gains in both constraint satisfaction and excess size relative to control.
Finally, ablations at T=15 show that a comment must carry outcomes. Comment-shaped noise lands within the noise of the no-comment arm, and a narrative of attempts without their outcomes is our worst arm at +70.0%, handing successors an unvalidated premise to extend. Within the schema the two fields recording what happened carry the reduction: dropping the recurrence count alone costs 37% of it.
Extraneous, noisy instructions degrade compliance with the true, correct instructions, and comments recover most of that loss. Inverting WildIFEval the same way carries the test to real prompts, and to instruction-following rather than count. We convert every benchmark item (Dj, Cj) to a world as in Section 4.1, but instead of making the maintainer learn all Dj instructions we seed it with Dj − K true instructions and G noisy ones drawn uniformly from other items' sets Di≠j. WildIFEval's constraints are human-written prose and ship with no code verifiers, so we score them with an arm-blind LLM judge that sees one constraint and one response and nothing else: no prompt, no arm label, no history. Across 64 worlds, for K=1 and G=16, those noisy instructions cost 24.1pp of correctness on the true instructions already in the prompt, 65.6% at G=0 against 41.5% here (95% CI: [−33.4, −14.9]pp).
Comments lift satisfaction from 50.4% to 62.0% over 3 maintenance rounds (11.6pp, 95% CI: [5.1, 18.3]pp), a 23.1% relative gain, against an uncommented maintainer given the identical prompt. Ablations show again that the latent reasoning encoded in the comment drives the behavior: comment-shaped noise lands 2.7pp from the uncommented arm (95% CI: [−4.4, 10.1]pp, covering zero). Every contrast in this subsection is a rate under one judge, so we re-scored all 6,336 verdicts under a second: the criteria reproduce, and the effect that judge measures differs from the one quoted here by 3.8pp (7.8pp against 11.6pp; 95% CI: [−1.9, +9.6]pp), which does not exclude zero.
Under assignment, then, we see that prompt comments encoding latent reasoning both reduce instruction count to its optimal minimum cover and buy back instruction-following in real-world prompts.
Empirical studies of agentic context files establish that context files earn their keep and only grow. Chatlatanagulchai et al. characterize AGENTS.md-class files, which accumulate content and rarely shed it; Tafreshipour et al. track prompt evolution in repositories, where additions dominate every other edit type; Lulla et al. show they pay for themselves, since repositories carrying one finish agent tasks faster and in fewer tokens. We measure that growth over the lifetimes of individual instructions, fine enough to estimate a deletion hazard and identify what drives it.
Agent memory systems already implement forgetting, each against an observable staleness proxy. MemGPT evicts on context overflow; Mem0 deletes on detected contradiction, and its graph variant and Zep timestamp superseded entries invalid; FSFM catalogs mechanisms from passive decay to safety-triggered deletion. Authored instructions admit no such proxy: an instruction does not become outdated merely by being old or rarely triggered. Instead, we show that an instruction's latent reasoning should be the primary driver of whether it should be retained or not.
Organizational written rules are artifacts people maintain for decades and have been studied extensively. Lehman's laws attribute a program's growth to a changing environment, the demand-side account; Zhou treats rule change as a stochastic process whose rate depends on a rule's own age; Schulz shows rule birth rates falling as the population densifies; and March et al. assembles both into an account of rule births, revisions, and suspensions. We use their insights to characterize the unique dynamics in agentic programming beyond compositional effects and staleness and further decompose a novel latent reasoning recoverability factor ρ̄(a) that depends on age a.
Software engineering solved comparable challenges decades ago via syntactic structures and engineering conventions. Literate programming argued that a program is addressed to a human reader; McConnell makes commenting the why standard practice; LaToza et al. find recovering that rationale is developers' most serious problem; architecture decision records institutionalize recording it at decision time; and commit messages exist to carry it and routinely fail. We directly build on both the structural and semantic insights to drive the core pillars of our result, showing that comments encoding an instruction's latent reasoning cut excess size at parity constraint satisfaction.
However, code-comment fidelity is notoriously weak and entire lines of research are dedicated to their divergence. Comments and the code they describe co-evolve poorly; detecting and addressing comment–code inconsistency are their own ML-for-SWE tasks. Our ablations identify that low-information, poorly-structured and misleading rationale carries the same deleterious effects into agentic coding.
Continual learning is our closest cousin but primarily targets tasks with drifting objectives. Networks trained on a task sequence overwrite what earlier tasks taught them; rehearsal and regularization toward parameters that mattered before preserve what a shifting objective would destroy; L2P, DualPrompt, and CODA-Prompt move that burden into the input, learning fixed-size prompt pools a frozen backbone selects from, which bounds growth by construction. Although our task is stationary, we borrow their frameworks and find complementary insights driven by fixed resource constraints: fixed parameters driving catastrophic forgetting under shifting objectives there, intermittent failures driving catastrophic remembering under (approximately) fixed instruction counts here.
IFEval scores responses against verifiable constraints; FollowBench grades progressively added constraint levels and WildIFEval collects real requests carrying many at once, both documenting that following degrades as the required set grows. Those benchmarks price the constraints a response must satisfy, the harm our motivation rests on. We invert them to price extraneous instructions instead and to make the minimum cover observable, without which excess size is unmeasurable.
If English is the new code, why don't we have comments yet? Fundamentally, we show two things: (i) agentic prompts grow because an instruction's latent reasoning decays faster than the instruction, and (ii) saving that latent reasoning in prompt comments at write time reduces instruction count and wins back instruction-following. Neither half should surprise us. Software engineering identified and solved the same problem decades ago and continual learning is working on solving its dual — catastrophic forgetting — right now. Although both settings differ in detail, both offer lessons in principle; however, LLM agents started driving real work before we could adopt their lessons. We eagerly borrow from both now.
Our results open three novel lines of work. Coding agent developers can give prompts a comment syntax, so an instruction's latent rationale can reach the next maintainer; our prototype, while promising, is just a prototype. Maintainers who manage agentic coding prompts can identify best practices: while they may overlap substantially with software engineering best practices, there are almost certainly differences. Finally, researchers can look for what (i) beats it, since prompt maintenance is continual learning over text and nobody has built its analogue of rehearsal or regularization, and (ii) measures it, at representative constraint counts and for constraints that are not mechanically verifiable. Every field is borrowing from AI right now. We should borrow back just as readily. The answer to catastrophic remembering was forty years old and one field over, and the next one might be too.
Improvements for AI systems
Based on the paper, here are the specific improvements you can make to AI systems:
-
Improvement: Implement a structured comment field (e.g.,
# rationale:) in agentic context files (CLAUDE.md, AGENTS.md) that stores the latent reasoning behind each instruction at write time. -
What the improved system can do: When a maintainer adds an instruction, the system prompts them to record why it exists (the failure it prevents, the hypothesis, the observed outcome). This metadata is stripped before execution but retained for future maintainers. This reduces excess instruction growth by 99.3% (from +211.3% to +1.4%) and improves instruction-following by up to 23.1%.
-
Improvement: Build an agent that tracks instruction age and author count, and uses the paper's finding that deletion hazard decays with age (log-hazard slope −0.032/commit) and with multi-author edits (β = −0.021).
-
What the improved system can do: Proactively flag instructions older than a threshold (e.g., 50 commits) or in files with many authors, prompting the maintainer to re-verify their rationale. This prevents the
catastrophic remembering
where prompts grow unboundedly (+226% over lifetime). -
Improvement: Use the paper's inversion of IFEval to create a verifiable world where the optimal prompt (minimum cover) is known. Integrate this as a test harness for any agentic prompt.
-
What the improved system can do: Automatically measure excess size (D/D⋆ − 1) and correctness for any prompt. It can then suggest deletions that don't risk regressions, because it knows the true constraint set. This makes prompt maintenance measurable and auditable.
-
Improvement: Replace free-text comments with a structured schema:
instruction, rationale, hypothesis, recurrence count, last verified. The paper shows that dropping the recurrence count alone costs 37% of the benefit. -
What the improved system can do: When a task fails, the system updates the recurrence count and last verified timestamp. This gives maintainers an evidence-based signal for which instructions are still necessary, enabling deletion decisions in O(1) instead of O(2 D).
-
Improvement: Train a model that, given a prompt with comments, can generate a minimal instruction set by reasoning over the comments' outcomes (not just the instructions' text).
-
What the improved system can do: At each maintenance round, it can propose a new prompt that is −5.8% excess size (vs. +60.4% without comments) at parity constraint satisfaction. It can also handle
mass rewrites
better, because it knows which instructions were actually excess vs. which were load-bearing. -
Improvement: Track the estimated recoverability ρ(d, a) of each instruction's rationale, using age and author count as proxies. When ρ drops below a threshold (e.g., 0.5), the system triggers a review.
-
What the improved system can do: Prevents the
ratchet
effect where prompts grow faster after a rewrite (+4.9% per commit vs. +4.1% before). It can schedule rewrites only when truly necessary, not as the default deletion mechanism (which currently handles 77.3% of instruction deaths). -
Improvement: Borrow rehearsal and regularization techniques from continual learning, but apply them to prompt maintenance. For example, periodically
replay
old instructions' rationales to the maintainer. -
What the improved system can do: Maintains a stable prompt size over time, even as tasks drift. This is the inverse of catastrophic forgetting—here, the system prevents catastrophic remembering by rehearsing why instructions exist.
-
Improvement: Use the WildIFEval inversion to detect and remove noisy instructions (those drawn from unrelated tasks). The paper shows noisy instructions cost 24.1pp of correctness.
-
What the improved system can do: Automatically scores each instruction's contribution to constraint satisfaction and flags those with negative marginal value. Comments restore 11.6pp of the lost correctness (23.1% relative gain), so the system can prioritize which instructions to keep.
-
Improvement: Separate the executor-facing instructions from the maintainer-facing comments, as the paper's harness does. The executor sees only
Dt; the maintainer seesDtplus allrd. -
What the improved system can do: Prevents the executor from being distracted by rationale text (which would degrade instruction-following), while giving maintainers full context. This is analogous to code comments being stripped at compile time.
-
Improvement: Detect when a prompt is about to be mass-rewritten (losing ≥50% of instructions) and instead suggest targeted deletions based on comment outcomes.
-
What the improved system can do: Reduces the 77.3% of instruction deaths that occur via wholesale rewrites. It can identify which instructions are truly excess (via comments) and delete only those, preserving the rest. This prevents the post-rewrite growth rebound (91.5% recovery in 10 commits).
Abstract
Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression costs O(2 D) in a prompt of D instructions. We name the resulting divergence catastrophic remembering, the inverse of catastrophic forgetting around which continual learning is organized. First, we characterize this phenomenon across 247,694 instruction lifetimes in 1,867 repositories: agentic prompts grow without bound, more than tripling over their lifetime (+226%), gaining +4.9 net instructions every commit; further, the older an instruction gets, the less likely it is to be deleted (log-hazard-0.032/commit). Then, we show that prompt comments can halt the growth: inverting IFEval yields verifiable worlds whose optimal prompts are known, and there comments encoding latent reasoning remove 99.3% of excess instructions (+211.3% to +1.4%). Finally, applying the same inversion to WildIFEval, we show that prompt comments can improve real-world agentic instruction-following by up to 23.1%. If English is the new code, why don't we have comments yet?
Sources
- Agent READMEs: An Empirical Study of Context Files for Agentic Coding
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory
- FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
- WildIFEval: Instruction Following in the Wild
- On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents
- MemGPT: Towards LLMs as Operating Systems
- Deep Just-In-Time Inconsistency Detection Between Comments and Source Code
- Learning to Update Natural Language Comments Based on Code Changes
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- Prompting in the Wild: An Empirical Study of Prompt Evolution in Software Repositories
- What Makes a Good Commit Message?
- Instruction-Following Evaluation for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection