The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond
cs.CL, cs.CV
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/data-privacy-stack/presidio
Terminology
Sources
- Gemma 4 Technical Report
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Magistral
- No Intruder, no Validity: Evaluation Criteria for Privacy-Preserving Text Anonymization
- The Enron Corpus: Where the Email Bodies are Buried?
- gpt-oss-120b & gpt-oss-20b Model Card
- BRATsynthetic: Text De-identification using a Markov Chain Replacement Strategy for Surrogate Personal Identifying Information
- OpenAI Privacy Filter: A Cross-Lingual, Cross-Domain PII Evaluation Across 32 Benchmarks
- REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering