LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images
cs.CV, cs.CL
Submitted: 2026-09-02
Updated: 2026-09-02
Terminology
Sources
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
- ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
- Qwen3-VL Technical Report
- RedactBuster: Entity Type Recognition from Redacted Documents
- ProPILE: Probing Privacy Leakage in Large Language Models
- OCR-IDL: OCR Annotations for Industry Document Library Dataset
- DocVQA: A Dataset for VQA on Document Images
- LayoutLM: Pre-training of Text and Layout for Document Image Understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models