What a Deletion Certificate Covers, and Where It Expires: Auditable Removal from a Support-Vector Memory
cs.LG
Submitted: 2026-07-13
Updated: 2026-09-11
Comments: 17 pages. Revised certificate scope and numerical analysis, including minimum-norm tie-break experiments and readout-disagreement measurements
Code: https://github.com/VyLabs-AI/sv-attention
License: http://creativecommons.org/licenses/by/4.0/
The gist: A dense key--value cache gives an operator no way to verify a deletion: it does not say which stored entries currently contribute nothing to the output, and it offers no reference state that the
Terminology
Abstract
A dense key--value cache gives an operator no way to verify a deletion: it does not say which stored entries currently contribute nothing to the output, and it offers no reference state that the edited memory should match. We build a context memory whose entries carry explicit weights and ask what a deletion certificate over it can cover and where it expires. A one-class support-vector boundary fit around the keys of a context at test time supplies the coefficients used in the readout, so that each key is active (positive coefficient) or reserve (zero coefficient). Removing a reserve key leaves the readout unchanged without a re-solve, and deleting an active key with a decremental solver reaches the same state as re-solving on the remaining keys at the same coefficient cap. A three-key construction shows the limit of both properties: an inert key can acquire positive weight after one more token is admitted, so reserve status certifies the present solve and does not license permanent pruning. In 1,200 double-precision trials on Gaussian, near-duplicate, clinical (MIMIC-IV vitals), and learned keys, every trial reached the reference state, all but one through the maintained update, with median gate-score disagreement below 10-6 in every regime and a worst case of 1.1 times10-2 associated with numerically near-tied solutions and partition disagreement; maintained deletion ran 24 -- 223 times faster than re-solving. Declaring a minimum-norm tie-break as part of the reference, at a tolerance above the two solvers' disagreement in objective value, brings the worst readout disagreement below 2% of the readout range in every regime and leaves a gate-score disagreement of up to 4.2 times10-3. A deletion certificate should name its reference state and validity horizon. We show what a system must retain, or refuse to admit, to extend that horizon.
Sources
- Differentiable Convex Optimization Layers
- OptNet: Differentiable Optimization as a Layer in Neural Networks
- Titans: Learning to Memorize at Test Time
- Machine Unlearning
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- SnapKV: LLM Knows What You are Looking for Before Generation
- Uncovering mesa-optimization algorithms in Transformers
- MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
- Parallelizing Linear Transformers with the Delta Rule over Sequence Length
- VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
- H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks