Reconstruction of Personally Identifiable Information from Proprietary Data in Supervised Fine-Tuned Models
cs.CR, cs.CL, cs.LG
Submitted: 2026-05-12
Updated: 2026-08-25
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- CLIOPATRA: Extracting Private Information from LLM Insights
- Extracting alignment data in open models
- AlpaGasus: Training A Better Alpaca with Fewer Data
- Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture
- Large-scale online deanonymization with LLMs
- Position: Privacy Is Not Just Memorization!
- Unmasking the Reality of PII Masking Models: Performance Gaps and the Call for Accountability
- Gemini: A Family of Highly Capable Multimodal Models
- Magicoder: Empowering Code Generation with OSS-Instruct
- Exploring Memorization in Fine-tuned Language Models
- Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models
- Eliciting the Priors of Large Language Models using Iterated In-Context Learning
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs