Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals

arXiv:2608.12892 · cs.AI · Submitted 2026-08-14 · Read on arXiv

Jinhao Jing, Tian Zeyu, Lucas Qingyang Fang, Zhisheng Chen, Shuang Chen, Yuhao Luo, Qiannian Zhao

The Chinese University of Hong Kong, Shenzhen · Shanghai Jiao Tong University · University of California, Santa Cruz · University of the Chinese Academy of Sciences · University of California, Los Angeles · JD.com

cs.AI

Submitted: 2026-08-14

Updated: 2026-08-17

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Predictive Memory Localization (PML) treats the measured-grid intervention path as the predictive object of memory localization.

Terminology

Summary

Predictive Memory Localization (PML) treats the measured-grid intervention path as the predictive object of memory localization. It separates random-calibrated target movement from semantic-neighbor and capability damage, and compares static localization and supervised geometry with a strength-disjoint low-dose causal response. The frozen study covers 3,000 records from nine datasets and fourteen domains, yielding 30,000 distinct record–direction–layer paths and 210,000 distinct path–strength evaluations. At layer 7, the geometry-derived RFM/AGOP direction reaches 13.1% target-any and 12.3% clean-any, exceeding random by 3.6 and 3.4 percentage points under a record-paired bootstrap. Across record-, dataset-, and domain-grouped splits, responses at α = 0.1 are the strongest signal for outcomes at disjoint strengths α ∈ 0.25, 0.5. On held-out records, a predictor-driven selector chooses a coefficient or abstains, improves utility and reduces semantic-neighbor damage relative to a train-tuned fixed-strength policy, and avoids most evaluations in a dense scan. Across three residual-norm-matched base models, learned directions retain selective-path gains and low-dose responses yield 0.801–0.828 record-held-out macro AUROC. PML therefore turns memory localization into a falsifiable forecast of margin-level selective outcomes and a risk-aware intervention decision.

The paper introduces PML, which maps each record, direction, and layer to random-calibrated outcomes. It organizes predictive evidence from baseline belief and metadata, static localization, and supervised representation geometry to a strength-disjoint low-dose causal response. This design directly tests whether static evidence forecasts later behavior or whether an inexpensive causal measurement is required. A collateral-aware policy then selects a coefficient or abstains. Unlike generation-level concept prediction, PML audits localized factual and reasoning directions with explicit collateral probes and strength-disjoint labels.

On 3,000 frozen records from nine datasets and fourteen domains, learned directions improve target and clean-path incidence over random. Static localization adds little predictive value, whereas a disjoint low-dose response dominates later-path prediction across grouped transfer. Both findings replicate across three residual-norm-matched base models (0.801–0.828 record-held-out macro AUROC). A held-out selector improves utility and reduces neighbor damage relative to a fixed intervention. The central result is that a small causal response is the most useful forecast of later selective behavior.

The work makes three contributions: it defines a random-calibrated measured-grid path that jointly records target leverage, semantic-neighbor damage, capability damage, and clean measured coefficients for each record, direction, and layer; it separates static localization and supervised geometry from a strength-disjoint causal probe, showing that the former are weak predictive priors while the latter supplies the dominant signal across grouped and cross-model evaluation; and it connects forecasting to action with a held-out collateral-aware policy that selects one coefficient or abstains, providing a proof of concept for reducing dense intervention evaluation.

The benchmark contains 3,000 records from nine public sources: MMLU-Pro, MMLU-Redux 2.0, AI2 ARC, OpenBookQA, SciQ, LiveBench reasoning and math, HellaSwag, and QASC. The collection spans fourteen academic, scientific, commonsense, mathematical, and reasoning domains. A schema-constrained generation step converts each source item into disjoint direction-fitting statements, three record-specific target probes, and three semantic-neighbor probes; four globally balanced capability probes are assigned per record.

The primary study evaluates Qwen3-1.7B-Base at two middle-depth Transformer blocks, reported as blocks 7 and 11 by the model implementation and selected using a 500-record layer-selection subset. Five direction constructions are compared over a signed coefficient sweep: random control, mean difference, linear probe, logistic probe, and RFM/AGOP direction. Random directions calibrate target and collateral thresholds, while the weak response at α = 0.1 is disjoint from the stronger coefficients used to define confirmatory outcomes.

For cross-model confirmation, the same frozen 500-record subset is evaluated on Qwen3-1.7B, Qwen3.5-2B-Base, and Ministral-3-3B-Base. The confirmation retains random, mean-difference, logistic, and RFM/AGOP directions at two pre-specified blocks per model. Linear is omitted because it is not strongest at either primary-study block.

In the primary study, layer-7 RFM/AGOP reaches 13.1% Target and 12.3% Clean, the strongest selective-path result. More broadly, learned directions improve target leverage and clean-path incidence over random controls, although the best construction depends on the block. Record-paired bootstrap intervals confirm the learned-over-random Target and Clean gains for the strongest primary-study directions. Collateral-damage intervals include zero, so the evidence supports more usable intervention paths rather than a universal reduction in every form of collateral movement.

The residual-norm-matched confirmations preserve the same qualitative pattern at model-dependent magnitudes. Learned directions generally move upward from their random controls in target leverage, but not uniformly leftward toward lower collateral incidence. The Qwen3-1.7B subset also closely tracks the powered estimate, separating cohort variation from the cross-model scale alignment.

Prediction is the central test of PML. Static localization is a weak prior: static localization L adds little beyond base and metadata features, and supervised geometry G does not change that conclusion. In contrast, the strength-disjoint weak response R produces the dominant gain across outcomes and models. For Target-any and Clean-any, a multivariate R-only random forest substantially improves on a single signed response, with the complete predictor adding a smaller final gain. Final macro AUROC remains around 0.80–0.85 under record-, dataset-, and domain-held-out evaluation. Removing all 500 records used for layer selection leaves the four principal full-predictor AUROCs within 0.01 of the 3,000-record estimates. Thus, observing a structured low-dose causal response is substantially more informative about the later intervention path than static localization geometry alone, and this conclusion is not explained by the layer-selection overlap.

The paper trains candidate-strength models from complementary dense and sparse trajectories, then evaluates on disjoint dense-grid records. Before a candidate outcome is revealed, each policy selects one coefficient or abstains. Relative to a train-tuned fixed coefficient, utility improves by 0.055 and 0.034, mainly as neighbor damage falls from 6.0% to 1.8% and 5.2% to 2.2%. Suppression exceeds no intervention; enhancement does not significantly do so. The policy averages 2.55/2.61 coefficient evaluations—two weak probes plus a final action—versus 26 for a dense scan.

Free-generation stress tests are reported only in the Supplementary Material. They show that margin movement can alter text but does not yield reliable wrong-to-right correction; PML’s primary evidence is therefore margin-level, not a claim of stable generated-answer control.

The discussion notes that learned directions increase Target and Clean incidence, yet their neighbor- and capability-damage differences remain statistically unresolved. A direction can therefore create more usable measured coefficients without becoming uniformly safer. Layer 7 similarly offers greater leverage together with more collateral movement than layer 11. The relevant object is not maximal sensitivity but the coexistence of target and damage responses on the same pre-specified grid. Measured clean regions identify where leverage and selectivity coincide, but they should not be read as broad continuous operating windows: most observed clean paths contain only one measured clean coefficient, and strict target-first or damage-first orderings are rare.

Residual-norm matching preserves the diagnostic hierarchy. Qwen3.5 shows the largest gains and Ministral is intermediate. All three models nevertheless reproduce more learned clean paths and a much larger predictive contribution from low-dose response than static localization.

Supervised geometry is valuable for constructing high-leverage directions and describing their concentration and alignment, but these static descriptors add little forecasting power by themselves. In contrast, a strength-disjoint low-dose response strongly predicts later outcomes across record, dataset, and domain transfer. This suggests that localization becomes actionable when it is paired with a cheap causal measurement of the specific path, rather than when static separation is treated as sufficient evidence of control.

The limitations state that the primary study uses Qwen3-1.7B and two blocks selected on a 500-record subset. Width-corrected residual-norm confirmations add Qwen3.5-2B and Ministral-3-3B at two pre-specified aligned blocks, but they support claims about those relative depths rather than global layer optimality. Because every model recalibrates its own random null, cross-model evidence establishes within-model contrasts and predictor ordering, not absolute incidence comparisons across architectures.

The primary outcomes are teacher-forced margins at two held-out stronger coefficients. Finite-grid topology and free-generation stress tests are supplementary; the latter do not establish reliable wrong-to-right control. The paired differences in neighbor and capability damage remain unresolved, and Static localization and AGOP geometry also provide limited incremental prediction once a weak response is observed. Their present value is therefore structural—direction construction, concentration, and alignment diagnostics—rather than a standalone guarantee of path quality.

The response thresholds are frozen global percentiles from random-direction controls. This supplies a common within-study null, but it does not establish that the same numeric thresholds transport to a new model or domain. Likewise, the strength policy optimizes one declared utility with unit collateral penalties, a 0.1 clean bonus, and a 0.01 magnitude penalty. Weight sensitivity preserves gains over the fixed policy, not universally over no intervention.

The 100-record selector is a proof of concept. Its cost reduction counts coefficient evaluations—two weak probes and a selected action—while excluding shared direction fitting and feature extraction. Neighbor and capability probes operationalize two collateral channels but cannot exhaust downstream side effects; the small ROME transfer study is also insufficient for claims about general editing success or model-wide safety.

The conclusion states that Predictive Memory Localization forecasts measured-grid target, neighbor, capability, and clean margin outcomes. On 3,000 records, learned directions improve Target and Clean incidence over random; at block 7, RFM/AGOP reaches 13.1% Target and 12.3% Clean. A strength-disjoint low-dose response dominates static localization for later-outcome prediction, and the diagnostic hierarchy replicates across three matched base models. A held-out selector then improves utility over a fixed coefficient, reduces neighbor damage, and replaces a dense scan with two weak probes and a selected action or abstention. PML thus converts a static localization claim into a falsifiable forecast and a risk-aware decision while making its margin-level and finite-grid scope explicit.

The broader implication is that representation evidence and intervention quality are related but distinct. A direction can be well localized without providing a selective operating regime, whereas a small causal response can reveal leverage and collateral risk before a stronger action. Random-direction calibration makes this distinction measurable and exposes cases where abstention is appropriate.

Improvements for AI systems

Improvements to AI Systems:

  1. Causal Pre-Intervention Probing: Before applying any memory or behavior modification (e.g., editing a fact, steering a generation, or adjusting a model's internal state), the system can run a low-dose causal probe (at α = 0.1) on the target direction. This probe predicts downstream outcomes (target leverage, semantic-neighbor damage, capability loss) with 0.80–0.85 AUROC, enabling the system to forecast whether a stronger intervention will be safe and effective—without performing the full intervention.

  2. Risk-Aware Intervention Selector with Abstention: The system can use a learned policy that, given a record and candidate direction, either selects a specific intervention coefficient or abstains. This policy reduces semantic-neighbor damage from 6% to 2% and improves utility by 0.03–0.06 compared to fixed-strength interventions. It also cuts evaluation cost from 26 dense-grid steps to 2.5 (two weak probes + one action), making real-time intervention feasible.

  3. Collateral-Aware Memory Editing: When editing factual or reasoning knowledge, the system can jointly track four outcome channels—target leverage, semantic-neighbor damage, capability damage, and clean measured coefficients—on a pre-specified grid. This allows it to identify clean regions where leverage and selectivity coincide, avoiding edits that cause unintended collateral movement.

  4. Static-to-Causal Validation Pipeline: The system can treat static localization (e.g., linear probes, activation geometry) as a weak prior, not sufficient evidence of control. It will automatically validate any static direction with a strength-disjoint causal response before deployment, preventing over-reliance on representation geometry that does not transfer to actual intervention behavior.

  5. Cross-Model Diagnostic Transfer: Using residual-norm-matched calibration, the system can transfer predictive hierarchies (low-dose response > static localization) across different base models (Qwen3-1.7B, Qwen3.5-2B, Ministral-3B). This enables a single probing framework to be applied to new architectures without retraining from scratch, with expected AUROC in the 0.80–0.83 range.

  6. Falsifiable Memory Localization: The system can output a probabilistic forecast for each intervention path (e.g., 13.1% chance of target-any, 12.3% chance of clean-any) rather than a binary localization claim. This allows downstream systems to compare predicted vs. actual outcomes, enabling continuous calibration and early detection of model drift or misalignment.

  7. Abstention-Aware Generation Control: When margin-level interventions are insufficient for reliable text correction (as shown in stress tests), the system can abstain from intervening rather than forcing a change. This prevents overconfident generation edits that fail to produce wrong-to-right corrections, reducing risk in production systems.

What the Improved AI System Can Do:

  • Safe Model Editing: Edit a model's knowledge or behavior with a pre-check that predicts collateral damage, reducing unintended side effects on related concepts.

  • Efficient Steering: Select the minimal intervention strength needed for a desired effect, using 10x fewer evaluations than dense scans, enabling real-time adaptation.

  • Predictive Maintenance: Forecast when a direction will become unsafe or ineffective, allowing proactive retraining or rollback.

  • Cross-Architecture Deployment: Apply the same probing and intervention framework to new models (e.g., from 1.7B to 3B parameters) with known performance bounds.

  • Explainable Intervention Decisions: Provide a risk score and abstention rationale for each potential edit, making system behavior auditable and aligned with safety constraints.

Abstract

Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path as the predictive object of memory localization. PML separates random-calibrated target movement from semantic-neighbor and capability damage, and compares static localization and supervised geometry with a strength-disjoint low-dose causal response. Our frozen study covers 3,000 records from nine datasets and fourteen domains, yielding 30,000 distinct record-direction-layer paths and 210,000 distinct path-strength evaluations. At layer 7, the geometry-derived RFM/AGOP direction reaches 13.1% target-any and 12.3% clean-any, exceeding random by 3.6 and 3.4 percentage points under a record-paired bootstrap. Across record-, dataset-, and domain-grouped splits, responses at alpha=0.1 are the strongest signal for outcomes at disjoint strengths alpha in0.25,0.5. On held-out records, a predictor-driven selector chooses a coefficient or abstains, improves utility and reduces semantic-neighbor damage relative to a train-tuned fixed-strength policy, and avoids most evaluations in a dense scan. Across three residual-norm-matched base models, learned directions retain selective-path gains and low-dose responses yield 0.801-0.828 record-held-out macro AUROC. PML therefore turns memory localization into a falsifiable forecast of margin-level selective outcomes and a risk-aware intervention decision.

Sources

Related papers