Blast Radius

arXiv:2608.07440 · cs.AI · Submitted 2026-08-07 · Read on arXiv

Algorithm Reconnaissance Division · Mankind Research Labs · North-West University · University of Pretoria

cs.AI

Submitted: 2026-08-07

Updated: 2026-08-27

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 58/100

The gist: "Blast Radius" by M.Y.

Terminology

Summary

Blast Radius by M.Y. Pitsane and Hope Mogale (Mankind Research Labs, Sandton; North-West University and University of Pretoria, RSA; arXiv:2608.07440v1 [cs.AI], 7 Aug 2026) addresses the growing problem of affordability and wasted tokens in agentic coding. The paper introduces Blast Radius, a predictive memory-management layer that estimates an incoming prompt's reach through coupled context and code channels, using retention likelihood and churn-weighted dependency reach, alongside two supplementary mechanisms: NECROPHORESIS, which makes eviction reversible by archiving dead context verbatim and replacing it with a compact skeleton, and Recurring Dead Matter (RDM), which identifies near-identical transcripts repeatedly injected by routine coding loops and buries recurrence classes using Laplace's rule of succession.

The paper reports that across seven OpenAI models (gpt-4.1 through the gpt-5.6 family), Blast Radius "managed to effectively reduce token consumption by 17–26%, achieved the lowest overflow rate among tested policies, and remained byte-exact reversible. Moreover, of 450 buried bodies, 378 were recurring dead matter and zero were recalled. The work is positioned as largely a work in progress toward the broader goal of Algosophy: making large language models and agentic coding more reusable and sustainable."

The motivating problem is that "a large language model driving a repository does not accumulate understanding the way an engineer does; it accumulates tokens. Each turn of an agentic coding loop re-submits the entire conversation, meaning the system prompt, the tool schemas, every prior file read, every diff, and every stack trace, so that the marginal cost of a turn grows with the integral of everything that came before it. The paper observes that the majority of that integral is dead: a file that was dumped in full to fix a typo on turn 3 is still occupying two thousand tokens on turn 40, long after the typo, the file, and the sub-task that motivated it have been concluded. In information-theoretic terms, dead context is pure epistemic entropy: it occupies bandwidth that could carry signal but carries noise instead."

The paper criticizes naive responses: "Sliding-window truncation drops the oldest tokens regardless of whether they were load-bearing; summarization memory compresses context through a second, itself-fallible model call and cannot be undone if it compresses something that mattered. Both trade a recoverable cost (tokens) for an unrecoverable one (information). Blast Radius's alternative is to estimate, before a turn runs, how far it will reach, both how much new context it will retain and how much structural surface its edits will touch, then use that estimate to license a forgetting operation whose downside is bounded by construction because it is exactly reversible."

The term blast radius is literal and two-channelled:

  • Context channel: the blast radius is the predicted increment to the retained working set; it sets the size of the eviction we must perform to keep the window from overflowing, and it selects which dead bodies to sweep.

  • Code channel: the blast radius is the churn-weighted set of files and symbols the turn's edits reach through the dependency graph; it is the impact surface that should trigger a checkpoint before it grows unreviewably large.

Both answer the question what is causally coupled to what I am about to do?—one over context tokens (temporal reach) and the other over the code graph (structural reach).

A session proceeds in turns t = 1, 2, …; the working context is a finite ordered set of bodies Ct = b1, …, bnt, where a body is a message-granular unit (a user prompt, an assistant mission, a tool transcript). Each body carries a token cost τ(b), an arrival index a(b), and a role. The load is Lt = Σ τ(b), constrained by the context window W. The paper separates two costs: Its carry cost is τ(b) per turn it remains resident; its information value ι(b) ∈ [0, 1] is the extent to which b still constrains a correct next action. The pathology: carry cost is constant in residency while information value decays: a concluded mission's ι falls toward zero, but its τ does not.

Liveness is defined probabilistically: Let qt(b) = P[b is live Ct] be its resurrection probability and dt(b) = 1 − qt(b) its death score. The death score is not observable at eviction time; it is a prediction.

Blast radius as reach: The context blast radius B(pt, Ct) is the expected increment to the retained working load attributable to executing the turn, with an estimator B̂ predicting B from prompt features such as length, declared tool intent, and the number and size of referenced files.

Eviction budget: Given a safety margin γ, "Ht = max(0, Lt + B̂(pt, Ct) − (1 − γ)W). If Ht = 0 the window has slack and no forgetting is required. If Ht > 0 we must reclaim at least Ht tokens before the turn runs."

The reversible sweep operator: Reclamation is performed by burial, not deletion. The midden M is an append-only, on-device archive. Each buried body gets a scent skeleton skel(b): a compact resident stand-in that names the mission, lists the files it touched, and carries the archival key, at a fixed cost τ(skel(b)) = σ tokens with σ ≪ τ(b). The sweep operator ΦS replaces S ⊆ Ct with skeletons and archives bodies to M; its inverse exhume replaces a resident skeleton by its archived body in a single keyed read. Net reclaimed mass is Δ(S) = Σ b∈S(τ(b) − σ).

Midden axioms (stated as assumptions): "(A2, reversibility) Archival is byte-exact… (A5, auditability) Every bury and exhume is appended to a ledger. (A7, redaction) Recognizable secrets are redacted before archival… (bounded exhumation) An exhumation costs at most κ tokens of overhead and O(1) reads."

Optimal sweep: Choosing S is a covering problem: reclaim at least Ht tokens at least expected regret. The optimal sweep is S⋆ = argmin Σ b∈S qt(b)κ subject to Σ b∈S(τ(b) − σ) ≥ Ht, over the HCRC-licensed candidate set Dt. This min-cost knapsack cover is NP-hard in general, but admits the standard greedy (1 + ln)-style approximation by descending efficiency e(b) = (τ(b) − σ)/(qt(b)κ + ε).

Recurring Dead Matter (RDM): Production telemetry revealed that "an agentic session is not a stream of unique missions; it is a loop. The agent greps the same symbols, reruns the same test suite, rebuilds the project from the first line to the newest… and checks the same version control status, turn after turn. Each of these routine calls injects a transcript into context that is near-identical to the one before it, and each such transcript dies the moment the next one arrives. The paper defines a recurrence class via a normalization map Σ that strips volatile content (counters, timings, hashes) from a transcript and retains its generator and stable head. A body is RDM if its class contains a strictly newer member. The newest member of each class is presumed live… every older member is RDM."

Resurrection probability for RDM is estimated from the class ledger by Laplace's rule of succession. Proposition 4.8 establishes that after kc burials of which ec were exhumed, q̂c = (ec + 1)/(kc + 2). "In particular, for a routine class that has never been exhumed (ec = 0), q̂c = 1/(kc + 2) decays monotonically with every recurrence, so a class that keeps dying and never resurrects earns immediate burial: the threshold test is passed after finitely many recurrences and stays passed. The resulting policy: keep the newest instance of every recurrence class resident, bury every older instance on sight (no threshold, no census wait), and let the ledger's own counts justify the aggression."

The code channel reads the abstract syntax tree the editor already parses and the dependency DAG the executor already walks (the Coefficient DAG of [1]), and scores reach over them. Given a dependency graph G = (V, E) and a seed set S0 of touched files/symbols:

  • Impact reach: Rk(S0) = v ∈ V: dG(S0, v) ≤ k, with each v carrying realized churn w(v) = added(v) + removed(v).

  • Radar encoding: The operator-facing rendering is a radar disk on which each touched file is a blip whose radial coordinate encodes its churn. The deployed encoding is r(v) = min(rmax, a + c√w(v)), a = 22, c = 3.1, rmax = 126. The square root is intentional: "taking r ∝ √w makes area linear in churn, so that perceived magnitude is proportional to realized change. This is an area-proportional encoding in the tradition of graduated-symbol cartography, corrected for the psychophysical tendency to under-read area."

  • Commit pressure: Πt = 1[∃v: tier(v) ≥ RISK], which fires when any file's accumulated churn breaches the risk tier, prompting a checkpoint before the reviewable surface grows too large. Deployed tiers are (θMED, θHIGH, θRISK, θDEADLY) = (50, 200, 500, 1000) lines.

The paper unifies both channels under a single mathematical substrate, a Polish space of context objects. The space has product decomposition X = Xprompt × Xcode × Xdep × Xarchive, with each item carrying a semantic representation, dependency neighborhood, retention likelihood r(x), churn estimate c(x), and recurrence score q(x). The composite metric is d(x, y) = αdsem + βdlex + γddep + δdtemp, which is what gives the space useful geometric structure: it determines what 'near', 'recurrent', 'dead', and 'reachable' actually mean.

Blast radius is defined as a measurable function B(x) = r(x)Σ y∈N(x) w(x, y) + λc(x), where the context-channel term… is the predicted retained-token increment; the code-channel term λc(x) is the churn-weighted structural reach.

NECROPHORESIS is formalized as a measurable map E: Xactive → Xarchive × S, with Theorem 6.4 proving reversibility: "E−1(a(x), s(x)) = x for all x ∈ Xactive. That is, E is a bijection onto its image, and exhumation is the application of E−1.… NECROPHORESIS does not destroy low-value context. It moves context from the active region of X into a reversible archive while retaining a compact sufficient skeleton for future retrieval."

Recurrence classes are formalized as closed balls [x]ε = y ∈ X: d(x, y) ≤ ε; the deployed normalization map Σ induces the metric d on the recurrence subspace: two bodies b, b′ are recurrences iff Σ(b) = Σ(b′), which corresponds to d(b, b′) = 0 on the quotient space.

The domination threshold (Corollary 6.7): burying x dominates carrying it whenever q(x) < (τ(x) − σ)m̄/κ, where m̄ is the expected number of remaining turns and σ is the skeleton cost.

The paper's central claim is Theorem 7.1 (Bounded downside of reversible sweep): for a body buried at turn t and required again m turns later (m = ∞ if never), "net saving(b) = (τ(b) − σ)·m′ − κ·1[m < ∞], m′ = min(m, turns remaining). The proof shows that the worst-case excess cost of burying b is exactly κ tokens, independent of m, whereas the saving grows linearly in the burial duration m′. This is the whole argument: Lossy forgetting (truncation, summarization) has an unbounded downside: if it discards something later needed, the information cannot be recovered and the turn may fail outright. Reversible forgetting caps that downside at a single small constant κ."

Corollary 7.2 (Domination threshold): "Burying b has non-negative expected token value iff q(b) ≤ (τ(b) − σ)m̄/κ. Since τ(b) ≫ σ and m̄ ≥ 1 while κ is a small constant, the right-hand side typically exceeds 1, in which case burial has positive expected value for every q(b) ∈ [0, 1], so carrying b is dominated."

Remark 7.3 stresses the threshold is measured: the deployed midden ledgers every burial and every exhumation, so the realized resurrection rate, the fraction of buried bodies ever exhumed, is a direct running estimate of E[q].… The framework thus closes its own loop.

An information-theoretic reading follows: under Assumption 7.4 that the HCRC-licensed candidate set carries near-zero information about the correct next action: I(Dt; A⋆t+1 Ct Dt) ≤ ε, Proposition 7.5 shows sweeping reduces resident token cost by Δ(S) while reducing the retained task-relevant information by at most ε. Reversibility further means any lost ε is recoverable on demand, so the loss is transient rather than structural.

Corollary 7.6 gives the monetary image: saved = πΣ(τ(b) − σ)m′b − πκN exhume, the area under the 'avoided carried tokens' curve, less a bounded exhumation tax.

The shipped policy is deliberately simpler than the general framework. The census runs each idle cycle over resident assistant bodies; a body is a candidate iff: "(i) it is an assistant mission, not a user turn; (ii) it lies outside the protected window of the last K = 3 missions, which are presumed live (d = 0); (iii) it is not already buried or marked immune; and (iv) its content exceeds a candidacy floor of 800 characters. Constants: τ(b) ≈ 41 × b chars, σ ≈ 60 tokens, recommended sweep threshold θ = 4000 tokens. Notably, The engine recommends; the operator's hand consents; only then does burial occur."

The midden is an on-device SQLite store under /.chalk with a bodies table (verbatim content, token count, scent, exhumation counter) and a ledger table (every bury/exhume with marker and timestamp). Redaction runs a fixed battery of secret patterns (provider keys, cloud credentials, token/password assignments) and truncates any match to a short prefix before archival. Burial is transactional: either all bodies in a sweep are archived and ledgered or none are.

The radar "renders every file the session wrote or edited as a blip… and raises the commit-pressure flag Πt when any file breaches the RISK tier at 500 churned lines. It is an ambient signal, not a gate: it advises the operator to checkpoint, and never blocks an edit."

The evaluation preregisters three research questions (RQ1: token cost reduction without degrading task success; RQ2: accuracy of the blast-radius estimate B̂; RQ3: realized exhumation rate vs. the domination threshold), run over five conditions: "(A) Carry-all: full history every turn, the baseline. (B) Truncation: fixed sliding window… (C) Summarization: MemGPT-style compaction… (D) Blast-Radius NECROPHORESIS (ours, deployed): census-gated reversible burial of concluded missions, exhumation on demand. (E) Blast-Radius + RDM (ours, full): condition D plus recurring-dead-matter reclassification. Each episode includes routine load of a real session: after every turn the agent rebuilds the project from zero, reruns the full test suite, and checks version control… conditions A–D must carry or lossily discard it, condition E reclassifies it. Models span seven OpenAI models spanning gpt-4.1 through the gpt-5.6 family (sol, luna, terra)."

Primary results (Table 2, median per-episode tokens and cost): Carry-all 43,053 tokens/ 0.055 with 5.21 overflows per episode; Truncation 34,662/ 0.044 with 5.07 overflows; Summarization 35,475/ 0.045 with 4.07 overflows; Blast-Radius (D) 39,568/ 0.051 with 5.14 overflows; Blast-Radius + RDM (E) 34,518/0.044 with 4.00 overflows. Success was 100% across all conditions.

Three preregistered predictions were confirmed: (P1) "the deployed census alone (D) cut median tokens from 43,053 to 39,568, a real but modest 8%, because mission burial cannot touch the recurring transcripts that dominate the load. Adding RDM (E) cut the median to 34,518, a 20% reduction overall and 17–26% per model across the gpt-4.1–gpt-5.6 ladder, at identical 100% success. The full policy matches the token economy of lossy truncation (34,662) while remaining byte-exact reversible, and posts the lowest overflow rate of any condition (4.00 versus 5.21 for carry-all)." (P2) success was 100% everywhere because each task is answerable from recent context, so the discriminating signal was cost and overflow rather than correctness. (P3) "Held decisively: across 68 mission burials under D and 450 burials under E, of which 378 were recurring dead matter buried on sight with no threshold wait, zero bodies were ever exhumed (Table 3). The aggressive rule-of-succession policy was empirically as safe as the conservative census."

Figure 7 shows the source of reclaimed mass under condition E: Routine traffic, the same build log, test rerun, and version-control chatter dying every turn, contributes 39% of the reclaimed mass and all of the improvement from D to E.

The paper concludes that "reversibility turns forgetting from a lossy gamble into a bounded downside bet, making provably dead context dominated whenever its resurrection probability falls below an explicit threshold that the system already measures. Blast Radius is the scoping layer beneath HCRC's gate licensed compaction, which dictates that verification decides what may be forgotten, while Blast Radius decides how much to forget, what to bury, and how far the next prompt can reach, while remaining ready to remember it again. The headline empirical result: 450 bodies buried, 378 recurring dead matter, zero exhumations. The work remains a work in progress toward the broader goal of Algosophy, making large language models and agentic coding more reusable and sustainable."

Improvements for AI systems

  1. Reversible context eviction with bounded downside.

Replace lossy sliding-window truncation and summarization with an archive-based burial system. Every evicted context unit is stored byte-exact on device, replaced by a compact skeleton, and recoverable via a single keyed read. The improved AI system can forget aggressively without risking unrecoverable information loss: the worst-case cost of forgetting something later needed is a small fixed exhumation overhead, not a failed turn.

  1. Predicted blast-radius estimation before each turn.

Before executing a turn, estimate (a) the expected increase in retained context load and (b) the churn-weighted set of files/symbols the turn will touch through the dependency graph. Use that estimate to compute an eviction budget and trigger checkpoints before the reviewable surface grows too large. The improved AI system can preemptively reclaim exactly enough tokens to avoid overflow, rather than reacting after the window fills.

  1. Recurring-dead-matter detection and immediate burial.

Detect near-identical transcripts repeatedly injected by routine loops—rebuild logs, test reruns, version-control status, repeated greps—by normalizing away volatile content and grouping by generator/stable head. Keep only the newest instance resident; bury all older instances on sight. The improved AI system can eliminate the dominant source of context waste in agentic coding loops: routine traffic that dies the moment the next identical transcript arrives.

  1. Laplace-rule resurrection scoring and self-measuring safety.

Maintain a per-class ledger of burials and exhumations, and estimate each class’s resurrection probability as (exhumations + 1) / (burials + 2). Since routine classes are never exhumed, their estimated resurrection probability decays monotonically, licensing increasingly aggressive burial. The improved AI system can justify its own forgetting policy from observed behavior, not from hand-tuned thresholds, and can keep burying the same routine class with confidence because the ledger proves it never resurrects.

  1. Churn-weighted dependency radar and commit-pressure gating.

Parse the AST and dependency DAG already available to the editor/executor, score reach over the graph, and render each touched file as a blip whose area is proportional to actual churn (using r ∝ √w so area is linear in churn). Raise a commit-pressure flag when accumulated churn breaches a risk tier. The improved AI system can advise operators to checkpoint before a change becomes unreviewably large, while never blocking edits.

  1. Minimum-regret sweep selection.

When eviction is needed, select which bodies to bury by solving a knapsack-cover problem: maximize reclaimed tokens per unit of expected resurrection cost, using a greedy approximation. The improved AI system reclaims the required token budget while minimizing expected future exhumation cost, instead of dropping the oldest tokens regardless of importance.

  1. Domination-threshold-based forgetting decisions.

Bury any context body when its estimated resurrection probability is below (token savings × expected remaining turns) / exhumation cost. Since token savings are large and exhumation cost is small, this threshold typically exceeds 1, meaning carrying the body is dominated by burying it for essentially any resurrection probability. The improved AI system can treat nearly all dead context as strictly better forgotten, with expected token savings linear in burial duration and downside capped at a constant.

  1. Monetary cost optimization for token-limited agents.

Use the realized savings formula: savings = price × (token cost − skeleton cost) × turns buried minus price × exhumation overhead per exhumation. The improved AI system can directly minimize API spend over long agentic sessions, cutting token consumption by 17–26% across model families while maintaining 100% task success and the lowest overflow rate among tested policies.

  1. Safe handling of routine-load-dominated agentic sessions.

Distinguish between concluded mission context and recurring-dead-matter context. The census alone only reduces tokens 8% because it cannot touch routine traffic; adding RDM reclassification reduces tokens 20% overall. The improved AI system can reclaim mass from build logs, test reruns, and version-control chatter—39% of reclaimed mass in the evaluation—without any exhumation ever being needed.

  1. Transactional, auditable, redacted archival.

Perform burial and exhumation as atomic transactions, append every operation to a ledger, and redact recognizable secrets before archival. The improved AI system can offer byte-exact reversibility, audit trails for every forgetting decision, and safe handling of credentials in archived context, making aggressive memory management acceptable in production environments.

Sources

Related papers