Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
cs.CR, cs.AI
Submitted: 2025-10-07
Updated: 2026-09-22
Journal ref: ICML 2026 - Forty-Third International Conference on Machine Learning, Jul 2026, Seoul, South Korea
Code: https://github.com/openlm-research/open_llama
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Scaling Instruction-Finetuned Language Models
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Mistral 7B
- 8-bit Optimizers via Block-wise Quantization
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Watermarking Text Data on Large Language Models for Dataset Copyright
- GPT-4o System Card
- GPT-4 Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs