MAPLE: Metadata Augmented Private Language Evolution
cs.CL, cs.AI, cs.CR, cs.LG
Submitted: 2026-02-26
Updated: 2026-09-17
Comments: COLM 2026
Code: https://github.com/elichien-google/MAPLE
License: http://creativecommons.org/licenses/by/4.0/
The gist: Differentially private (DP) fine-tuning of large language models (LLMs) requires massive compute and full model access, which rules out state-of-the-art proprietary APIs for general users.
Terminology
Abstract
Differentially private (DP) fine-tuning of large language models (LLMs) requires massive compute and full model access, which rules out state-of-the-art proprietary APIs for general users. Generating DP synthetic data offers a practical workaround. This approach also allows for transparent exploratory data analysis and arbitrary reuse across downstream tasks, sidestepping the rigid constraints of a model's parameter space. Private Evolution (PE) provides a promising API-based framework for generating this data, but its success relies heavily on initialization. If the private data distribution falls too far outside the foundation model's pre-training priors -- a common issue in highly specialized domain -- PE struggles to align with the target data. This misalignment causes poor convergence, degraded utility, and wasted API calls. To solve this initialization bottleneck, we introduce Metadata Augmented Private Language Evolution (MAPLE). MAPLE extracts DP tabular metadata and uses in-context learning to firmly ground the initial synthetic distribution in the target domain. Our evaluations on domain-specific text generation tasks show that MAPLE yields a strictly better privacy-utility trade-off, converges significantly faster, and sharply reduces API costs compared to baseline PE methods.
Sources
- The Llama 3 Herd of Models
- ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control
- Harnessing large-language models to generate private synthetic text
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy
- VaultGemma: A Differentially Private Gemma Model
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Synthesizing and Adapting Error Correction Data for Mobile Large Language Model Applications
- Contrastive Private Data Synthesis via Weighted Multi-PLM Fusion
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering