Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension
Tingan Jin, Shuhang Dong, Haosong Li, Chung-Hsien Chou
UCLA · Independent Researcher · Cal Poly Pomona
stat.ML, cs.CV, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 15 pages, 3 figures. Code and results included as ancillary files. Tingan Jin, Shuhang Dong, and Haosong Li contributed equally
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper argues that iterative erasure count is not an affine-invariant concept dimension.
Terminology
Summary
This paper argues that iterative erasure count is not an affine-invariant concept dimension. The authors show that both the stopping count and cumulative removed rank from iterative null-space projection can change under an invertible reparameterization that preserves all information, so neither quantity is intrinsically a concept dimension.
The paper distinguishes model-defined population quantities—generating dimension, sufficient linear dimension, and minimum guarding rank—from procedure-defined quantities such as stopping count and cumulative edit rank. The authors emphasize: Iterative erasure therefore returns a procedure-relative estimand jointly determined by representation geometry and the full measurement procedure, not a semantic dimension by itself.
Proposition 1 (Coordinate dependence): For d ≥ 2 and nonzero w, there exists an invertible A for which the mapped-back Euclidean edit differs from the Euclidean edit in the original coordinates, although xw = (xA)(A−1w) for every x. This shows that an invertible reparameterization preserves all pre-edit linear predictions but changes the erased subspace and every later refitted direction.
Proposition 2 (Complete transported-metric affine equivariance): For any positive-definite coefficient metric GX fixed from the intact representation, if the metric is transported as GZ = ATGXA, the probe is equivariant (BkZ = A−1Bk), regularization is transported under A, and tie-breaking is resolved equivariantly, then for every k, ZTkZ = (XTk)A. Scores, cumulative edit rank, and every score-based stopping count are identical in both parameterizations. Exact covariance is a natural corollary, but the paper stresses: Proposition 2 does not privilege covariance as a semantic geometry; any fixed positive-definite metric transported by the same tensor law obeys the result.
Theorem 1 (Cumulative-QR ridge count non-identification): In a population Gaussian construction with S, N N(0,1), L = sign(S), and Xa = (S + aN, N), the iteration count and cumulative edit rank are one for a = 0 and two for every a ≠ 0, although every Xa has sufficient linear dimension one and minimum linear guarding rank one. This holds for ridge λ ∈ 0,.1, 1, 10 and for both isotropic and transported penalties.
Theorem 2 (Full-QR multivariate count non-identification): For a two-output construction with Xa = (S + aN, N) ∈ R2ʳ, Y = S, the cumulative edit rank is r for a = 0 and 2r (the ambient dimension) for every a ≠ 0, despite common sufficient dimension and minimum guarding rank of r. For the motivating video analysis with r = 2, this changes the reported QR direction count from two to four.
The authors implement the motivating finite procedure on a controlled angular target with θ Unif[−π, π], S = (sin θ, cos θ), N N(0,.5I2), and Xa = (S + aN, N). Following the published specification, they use Adam/MSE/QR with 100 epochs, learning rate 10−3, weight decay 10−4, and an 80/20 split. Results at n = 4,000:
-
Identity mixing (a = 0) stops after one accepted update in all 20 runs
-
Each tested shear a ∈.5,.75, 1, 1.25, 2 accepts at least two updates in all 20 runs
-
At a = 1, post-first-edit held-out R2 is.195 and circular MAE is 59.5° on average; capped mean K is 5.75 (range 3–8), with two of 20 runs right-censored at eight
-
The decisive identity/a = 1 separation holds in all ten runs at minibatch sizes 64 and 256
The paper notes: Under the literal sequential recurrence, rank(I − TK) has reached four while rank(TK) remains about two, so 2K is not an independent removed-subspace dimension.
The visual study uses hand–object contact in frozen V-JEPA2 and DINOv2 representations. The authors stress: "Contact is not offered as a pure semantic primitive: 100DOH contains object-configuration shortcuts, and TouchMoment includes approach and action phase. Those limitations make it a demanding audit target rather than evidence for a universal contact mechanism."
Key results on V-JEPA2 layer 21 on 100DOH validation:
-
After mapped-coordinate restandardization, mean rank-ten AUROC is.599 at κ = 1 and.666/.673/.745/.791 at κ = 2/3/5/10, then.853/.840 at 100/1,000
-
Source-video-paired 95% bootstrap intervals for the mean mapped-minus-identity rank-ten difference are [.008,.134] already at κ = 2 and [.110,.275] at κ = 10
-
Unshrunk empirical-covariance intervention gives the identical complete trajectory for all maps (maximum prediction discrepancy 2.11 × 10−6), matching the corollary of Proposition 2
-
Twenty Haar orthogonal maps reproduce the complete Euclidean AUROC trajectory (maximum prediction discrepancy 4.27 × 10−6)
The fresh-attacker audit (official-training-only, source-disjoint A/B/C roles) shows: "Rank-zero AUROC is identical at.822. At rank ten, identity averages.665 across five source-disjoint partitions, whereas the five κ = 10 maps average.776 across 25 map–partition runs. The paired difference has mean +.111 and empirical range [.075,.155], and is positive for all 25 pairs."
The synthetic cross-fit calibration shows that in 1,024 dimensions, sample-estimated MP/SAL leaves independent AUROC.881,.754,.683,.622, and.572 at eraser sample sizes 100, 1,000, 2,000, 4,000, and 8,000, despite the known population rank-one guard. The population oracle gives.499 at every size. The paper concludes: Large residual access is therefore compatible with high-dimensional estimation error alone.
The authors explicitly acknowledge: Our affine theorem and latent construction concern linear representations and linear interventions. They do not identify nonlinear, tokenwise, spatial, temporal, or causal concept organization.
They also note that OAS and Ledoit–Wolf reduce covariance estimation problems but do not create a canonical semantic metric
and that Historical official-test exposure was extensive, no new untouched holdout is available, and none of those results is confirmatory.
The paper's central conclusion: "Many erased directions need not mean many semantic dimensions. Population ridge count changes from one to two for the same rank-one concept. Under full two-output QR removal, cumulative edit rank changes from two to the ambient dimension four, while a complementary sequential construction reaches ambient dimension eight."
The authors recommend: "Reports of iterative erasure count should specify the metric, probe loss, regularizer, covariance estimator, sample, orientation rule, stopping rule, and test-use policy, and should separate generating dimension, sufficient dimension, encoded cross-covariance rank, and guarding rank. Without those commitments, iterative erasure count is not an affine-invariant concept dimension."
Improvements for AI systems
Improvement 1: Metric-Transparent Concept Erasure
The AI system will explicitly track and report the coefficient metric (e.g., covariance, transported metric) used in any iterative erasure or concept-removal procedure. It will automatically flag when a reported concept dimension count
is procedure-relative rather than intrinsic, and will require the user to specify the metric, probe loss, regularizer, and stopping rule before producing a final count. This prevents misleading claims about semantic dimensionality in downstream applications like fairness auditing or model editing.
Improved capability: When a user asks to remove the concept of gender
from a model, the system will output not just the number of erased directions but also a full specification of the metric and procedure, and will warn if the count changes under an invertible reparameterization (e.g., if using a different covariance estimator). It can then generate a robustness report showing how the count varies across valid metrics.
Improvement 2: Reparameterization-Invariance Auditing
The AI system will automatically test whether any proposed concept-erasure result is invariant under invertible linear transformations of the representation space. Given a representation X and a target concept w, it will sample random invertible matrices A, apply the erasure procedure in the transformed space Z = XA, map results back, and compare the final erased representations and downstream predictions. If the procedure is not equivariant (as shown in Proposition 1), the system will flag the result as coordinate-dependent and refuse to treat it as a semantic ground truth.
Improvement 3: Dimension-Count Disambiguation
The AI system will separate four distinct quantities that are often conflated in iterative erasure: (a) generating dimension (intrinsic latent rank), (b) sufficient linear dimension (minimal rank for full prediction), (c) encoded cross-covariance rank (rank of the concept–feature covariance), and (d) guarding rank (minimal rank needed to block a specific attacker). When a user reports we found 4 concept dimensions via QR,
the system will automatically parse the method and output which of these four quantities is actually being estimated, and will provide confidence intervals based on the finite-sample behavior described in the paper.
Improvement 4: Sample-Size-Aware Erasure Reporting
The AI system will incorporate the paper's finding that high-dimensional estimation error alone can produce large residual access (AUROC.881 at n=100 vs.499 at population oracle). When reporting erasure effectiveness, it will automatically compute and display the expected residual access under a null hypothesis of no true concept, given the eraser sample size and ambient dimension. It will flag any result that falls within the estimation-error envelope as not evidence of residual concept information.
Improvement 5: Procedure-Equivariant Default Settings
The AI system will default to using a transported metric (GZ = ATGXA) whenever a reparameterization is applied, ensuring that iterative erasure counts and cumulative edit ranks remain invariant (per Proposition 2). It will also default to equivariant tie-breaking rules and will document any deviation from these settings. This makes the system's outputs reproducible and comparable across different representation spaces.
Improvement 6: Attacker-Role Separation in Audits
The AI system will implement the paper's fresh-attacker
audit protocol, which uses source-disjoint training/validation/test splits for the erasure procedure and the downstream attacker. It will automatically detect if the same data is used for both fitting the eraser and evaluating residual information, and will warn that any observed residual concept
may be due to overfitting rather than true semantic content.
Abstract
How many directions does a neural representation use to encode a concept? A common answer repeatedly erases probe directions and reports the stopping count or cumulative removed rank. We show that both quantities can change under an information-preserving invertible reparameterization, so neither is intrinsically a concept dimension. We distinguish model-defined population quantities (generating dimension, sufficient linear dimension, and minimum guarding rank) from procedure-defined quantities such as stopping count and cumulative edit rank. In a population Gaussian construction, an invertible shear preserves the prediction problem and all three quantities, yet changes the cumulative Euclidean erasure count from one to two. The separation holds for Moore--Penrose ordinary least squares and every finite nonnegative ridge weight. For a two-output full-QR procedure matching our motivating video analysis, cumulative edit rank similarly changes from two to the ambient dimension four. Conversely, the complete cumulative metric-QR trajectory is affine-equivariant when its positive-definite metric, probe, regularizer, and tie-breaking are transported consistently; exact covariance is one corollary, not a canonical semantic metric. In a known-rank finite-sample Adam/QR calibration, identity mixing stops after one accepted update in all 20 large-sample runs, whereas each tested shear a in.5,.75,1,1.25,2 accepts at least two updates in all 20 runs. Controlled reparameterizations of frozen V-JEPA2 features preserve rank-zero predictions yet alter later Euclidean trajectories under practical optimization. These visual contact experiments are stress tests, not estimates of contact dimension. Iterative erasure therefore returns a procedure-relative estimand jointly determined by representation geometry and the full measurement procedure, not a semantic dimension by itself.
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey