Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection
Zhen Yang, Mengqi Wang, Gengda Zhao, Mo Zhou, Jianwei Wang, Wenjie Zhang
The University of New South Wales
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 14 pages, 7 figures. The first two authors contributed equally
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: The paper proposes CALIBDCD, a calibration framework for feature-based data contamination detection (DCD) in large language models (LLMs).
Terminology
Summary
The paper proposes CALIBDCD, a calibration framework for feature-based data contamination detection (DCD) in large language models (LLMs). The problem addressed is that modern LLMs undergo post-training (instruction tuning, preference optimization, reasoning-oriented training), which alters model outputs and shifts membership features, reducing the separability between members (texts in the pre-training corpus) and non-members. The paper states: "Recent state-of-the-art DCD methods follow a feature-based paradigm that derives membership features from the input text and the corresponding model output. However, most modern LLMs undergo post-training, such as instruction tuning, preference optimization, and reasoning-oriented training, which can alter model outputs and shift the corresponding membership features, thereby reducing the separability between members and non-members."
CALIBDCD comprises two main components: (1) Multi-View Shift Detection, which evaluates controlled prompt variants on known non-member texts and consolidates the most informative views to identify recurring feature shifts,
and (2) Bounded Feature Correction, which selectively adjusts feature components aligned with the detected shifts and controls the correction extent to preserve useful detection information.
Multi-View Shift Detection proceeds in three stages: Multi-View Shift Measurement evaluates known non-members under multiple controlled query variants (eight universal response-format views and two assistant-generation-boundary views tailored to each target-model family) and measures feature changes relative to the original query. FPP-Based View Ranking prioritizes views using false-positive pressure (FPP), defined as the average positive score increase over the calibration non-members,
retaining the three highest-FPP views. Cross-View Shift Consensus estimates a score-guided shift subspace for each selected view via SVD, then combines view-specific subspaces through an average projector, retaining directions with cross-view support above a threshold τ.
Bounded Feature Correction constructs candidate corrections that attenuate, rather than completely remove, feature components aligned with the consensus shift directions. The transformation is Ar,λ = I − λBrB⊤r, where λ ∈ [0,1] controls attenuation strength. Controlled Correction Selection then selects the final correction based only on known non-members, maximizing the average reduction in non-member scores: J(r, λ) = (1/Dcal) Σ [q(zs,e0) − q(zs,e0 Ar,λ)].
Experiments evaluate CALIBDCD on four benchmarks (BookTection, BookMIA, ArxivTection, WikiMIA) and three target LLMs (Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, DeepSeek-R1-Distill-Qwen-7B), applied to two feature-based detectors (VeilProbe and DPDLLM). The paper reports: "CALIBDCD improves both AUC and TPR@5%FPR in all 24 settings, with average gains of 2.1% and 7.0%, respectively. The largest gains reach 7.0% in AUC for DPDLLM on BookMIA with DeepSeek and 15.0% in TPR@5%FPR for VeilProbe on BookTection with Qwen."
Ablation studies show that removing cross-view consensus produces the largest AUC reduction of 1.3% and decreases TPR@5%FPR by 3.9%, while replacing FPP-based selection with Random-3 Views reduces AUC by 0.6% and TPR by 3.1%. Fixed full correction (λ=1) reduces AUC by 0.9% and TPR by 2.6% across the 18 partial-correction settings. A decision-level recovery analysis found that 90 of the 201 identified false positives are recovered as true negatives, corresponding to a recovery rate of 44.8%.
Adaptive correction selection improves average AUC from 89.4% to 90.0% for VeilProbe and from 66.3% to 66.6% for DPDLLM, with TPR@5%FPR increasing from 60.3% to 63.9% and from 17.0% to 19.7%, respectively.
The paper concludes: "Across four benchmarks, three post-trained LLMs, and two detector families, CALIBDCD improves both AUC and TPR@5%FPR in all 24 settings, with average gains of 2.1% and 7.0% and maximum gains of 7.0% and 15.0%, respectively, without modifying the target LLM or detection-time query process." Limitations noted include that the method requires known non-members, may not address nonlinear or sample-specific false-positive sources, and that corrected shifts may reflect dataset artifacts or decoding behavior rather than post-training alone.
Improvements for AI systems
Improvements to AI Systems:
-
Post-Training-Aware Membership Inference: Enhance LLM-based data provenance systems by integrating CALIBDCD’s multi-view shift detection to automatically recalibrate membership features after any post-training step (e.g., RLHF, reasoning distillation). This allows AI systems to maintain accurate detection of copyrighted or sensitive training data even when model behavior changes, improving compliance auditing.
-
Adaptive False-Positive Suppression in Retrieval/Generation: Incorporate bounded feature correction into AI assistants that retrieve or generate text from corpora. The system can dynamically attenuate spurious membership signals (e.g., overconfidence on memorized phrases) without losing genuine relevance, reducing hallucinated attributions and improving factual grounding.
-
Calibration for Few-Shot and Instruction-Tuned Models: Use the FPP-based view ranking to select optimal prompt variants for calibration in production LLM APIs. This enables AI systems to self-calibrate on-the-fly with minimal non-member data, improving reliability of confidence scores for downstream decision-making (e.g., medical or legal Q&A).
-
Cross-View Consensus for Robust Anomaly Detection: Apply the consensus subspace estimation to any AI system that monitors model outputs for distribution shifts (e.g., drift detection in deployed chatbots). This improves early warning of adversarial or out-of-distribution inputs by identifying stable shift directions across multiple query formulations.
-
Controlled Correction for Bias Mitigation: Use the bounded attenuation (λ ∈ [0,1]) to selectively reduce feature components linked to systematic biases (e.g., demographic skew) identified via controlled prompt variants, while preserving task-relevant information. This yields fairer AI outputs without full debiasing that harms performance.
-
Benchmark-Agnostic Evaluation Harness: Build an automated calibration pipeline that, given any post-trained LLM and any feature-based detector, runs the 24-setting evaluation protocol to report AUC and TPR@5%FPR gains. This allows AI developers to continuously validate detector robustness after each model update.
What the Improved AI System Can Do:
-
Detect data contamination (e.g., leaked test sets) in post-trained LLMs with 2–7% higher AUC and 3–15% higher true-positive rate at low false-positive rates, across diverse domains (books, arXiv, Wikipedia).
-
Recover 44.8% of previously false-positive membership flags as true negatives, reducing unnecessary privacy alerts.
-
Self-calibrate using only a small set of known non-members, without altering the target LLM or query interface—enabling seamless integration into existing AI pipelines.
-
Provide interpretable shift subspaces (via SVD consensus) that explain why a text is flagged as member/non-member, aiding human oversight.
-
Maintain detection accuracy across instruction-tuned, preference-optimized, and reasoning-distilled model variants, ensuring long-term reliability as models evolve.
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Constitutional AI: Harmlessness from AI Feedback
- The Llama 3 Herd of Models
- Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering