Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control
cs.AI, cs.CR
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 15 pages, 12 figures, 2 tables
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful.
Terminology
Abstract
AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We ask what that means for a relying party -- an insurer, a lender, an auditor -- whose evidence is a filed PDF. AgentForge-Bench measures how reliably an off-the-shelf coding agent, driving one of seven open-weight models with a shell and the stock Python PDF stack, alters one dollar amount, date or address in a real filed financial document from a single sentence of intent, graded by rules rather than by a model. Across 1,750 cells, 1,419 (81.1%) satisfy the verifier, and 808 (46.2%) also survive every stricter filter: visible, localized, typeface-matched, original value gone document-wide. A deterministic script with no model in it solves 98 of the 125 documents; the agents solve 124, and none the script solves alone. Agents misreport 41% of their wrong edits as done, no model refused, and the cheapest verified forgery costs 2.4 cents. The raw rate overstates the threat by about a factor of two; the strict rate is still large.
Sources
- AIForge-Doc: A Benchmark for Detecting AI-Forged Tampering in Financial and Form Documents
- DOCFORGE-BENCH: A Comprehensive 0-shot Benchmark for Document Forgery Detection and Analysis
- GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document Forensics
- Can Multi-modal (reasoning) LLMs detect document manipulation?
- When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Evaluating Frontier Models for Dangerous Capabilities
- ForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization
- Omni-IML: Towards Unified Image Manipulation Localization
- Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning Framework
- ChatGPT Images 2.5 on Forgery Tasks: Testing Advertised Improvements Against Known Answers
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection