What Users See Is Not What Models Read: Split-View PDFs in Document-to-LLM Supply Chains

summary

Video file (mp4)

The gist

The core finding of this research is that document-to-LLM pipelines suffer from semantic integrity failures because PDF renderers and extractors operate independently, allowing attacker-controlled or

In short

Document-to-LLM pipelines fail because PDF rendering and text extraction happen separately, creating 'split-view PDFs.' This allows hidden, attacker-controlled text to be consumed by AI models while remaining invisible to the human user. The research identifies 25 specific extraction gaps across four categories, proving that deployment paths and ingestion stacks dictate which vulnerabilities are exposed.

Key concepts

Split-View PDF
A PDF where what a human sees (the rendered page) is different from what an extractor pulls out as text. This happens because the rendering process and the text extraction process operate independently, allowing malicious or unexpected data to exist only in the extracted stream, not on the visual page.
Extraction Gaps (EG)
Specific technical flaws within PDF specifications that allow for inconsistent text extraction. These gaps range from font-level overrides to reading-order splits. They are the root causes of split-view PDFs, enabling text to be hidden or reordered only for the downstream LLM.
Two-Tier Benchmark
A testing method used by researchers that separates testing into two stages. Corpus A tests if an extractor detects a specific gap using inert tokens, while Corpus B tests how these gaps affect actual LLM tasks like summarization or question answering.

Terminology used across episodes

This episode discusses

The paper

What Users See Is Not What Models Read: Split-View PDFs in Document-to-LLM Supply Chains · Read on arXiv

Tulane University

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "What Users See Is Not What Models Read".

Elias: The core finding of this research is that document-to-LLM pipelines suffer from semantic integrity failures because PDF renderers and extractors operate independently,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, we're looking at this paper, "What Users See Is Not What Models Read: Split-View PDFs in Document-to-LLM Supply Chains," and it seems the core idea is that there's a serious problem with how document-to-LLM pipelines handle PDFs. Essentially, the authors show that these pipelines have semantic integrity failures because the tools responsible for rendering a PDF page and extracting its text operate independently, which lets attackers sneak in or extract text that looks fine to a human but carries different meaning for the AI model <ref:2606.15020#pg0>.

Elias: That's what caught my attention too, Nadia; it suggests a fundamental split-view PDF where the rendered page presents one set of semantics while the extracted text carries something else entirely, creating a supply chain issue upstream before the model even gets its input <ref:2606.15020#pg0>. It makes you wonder what kind of text an AI might actually be reasoning over when it's fed this divergent information.

Priya: From my side, I’m interested in what this means for the data itself; if the extracted text is semantically different from what a user sees, then any subsequent analysis or summary generated by an AI could be based on misinformation that was hidden from view <ref:2606.15020#pg1>. It raises questions about the trustworthiness of LLM outputs derived from these documents.

Nadia: Exactly, Priya; the paper claims they've identified twenty-five distinct extraction gaps across four families—semantic overrides, hidden semantic injection, reading-order splits, and font-decoding splits—and that these gaps are often not seen in previous work <ref:2606.15020#pg1>. It really highlights how much the PDF specification itself allows for these representation gaps between rendering and extraction <ref:2606.15020#pg0>.

Elias: I agree, Nadia; focusing on those specific mechanism-level mismatches, like the font-level /ToUnicode CMap override or span-level /ActualText marked-content attribute overrides mentioned in the abstract, shows that this isn't just a bug in one tool but a tolerated gap in how PDFs are processed generally <ref:2606.15020#pg0>. It points to a deeper issue with the PDF model itself.

Priya: And when you look at those four families, particularly reading-order splits and font-decoding splits, it suggests that the divergence isn't always malicious injection but can stem from how different systems interpret the structure of the document versus how a human reads it <ref:2606.15020#pg0>. That distinction between spatial reading order and content stream order is something we need to consider for privacy measurement.

Paper summary: Nadia: That's a good point, Priya; and Elias, you mentioned the four families; are there any particular gaps that seem more exploitable or easier to trigger in practice based on what you’ve seen? We're trying to figure out if this is just theoretical or something we can actually see happening on the ground <ref:2606.15020#pg1>.

Elias: The paper does mention that fourteen of those twenty-five extraction gaps don't have an exact path or mechanism-level match in prior studies, which is significant because it means we’re looking at novel ways to exploit these PDF structures <ref:2606.15020#pg1>. The authors also built a two-tier benchmark to systematically test whether these gaps actually propagate into the final LLM output through their summary and QA attacks <ref:2606.15020#pg1>.

Priya: That two-tier benchmark sounds like it’s very rigorous for separating the parser layer issues from the downstream impact on the model, which is important for understanding what the data actually shows <ref:2606.15020#pg1>. It helps us see if a hidden text insertion in a document translates into a flawed answer from an AI service.

Nadia: Right, so it’s not just finding the gap; it’s proving that the gap actually causes divergence when you test it against sixteen different processing stacks and seven commercial LLM services <ref:2606.15020#pg0>. That scale of testing shows how widely this problem is exposed, showing coverage ranging from twelve to twenty-one out of the gaps for each service <ref:2606.15020#pg1>.

Elias: It confirms the idea that exposure isn't tied to one specific model identity but rather to the ingestion stack—the APIs, cloud backends, and local runtimes—which is a crucial detail for understanding where we need to apply our scrutiny <ref:2606.15020#pg1>. The paper also suggests that these canaries serve as a way to fingerprint which specific PDF parser or loader path an LLM application is likely using <ref:2606.15020#pg1>.

Priya: Fingerprinting the ingestion stack sounds like a strong diagnostic tool; if we can identify the parser being used, it tells us exactly which part of the supply chain to focus our efforts on securing for privacy and integrity <ref:2606.15020#pg1>. It moves the discussion from just "is it broken" to "which component is causing the breakage."

Paper summary: Nadia: And that leads directly into what these gaps imply for security; if we can fingerprint the path, we can potentially target defenses more effectively <ref:2606.15020#pg1>. But what about defense? The authors tested a static screening scanner that flags all twenty-five benchmark gaps in their self-test, looking for things like "extractor-side replacement text" or "non-painted text operators" <ref:2606.15020#pg1>.

Elias: That scanner is interesting because it’s designed to catch the structural conditions that lead to these issues, but the paper notes that existing safety filters, like those in OpenDataLoader, often fail because they are based on fixed rendering-mismatch heuristics rather than checking for actual semantic consistency <ref:2606.15020#pg2>. It sounds like they’re not a fix for the underlying representation issue.

Priya: So, it suggests that relying solely on existing filters isn't sufficient because those filters don't understand the semantic divergence between what is visually present and what is extracted <ref:2606.15020#pg2>. This reinforces the need for solutions that check for consistency across modalities rather than just blocking specific text patterns.

Nadia: And when we look at defenses, the paper points to vision-based processing as the strongest defense against these text-layer split-view attacks, provided it's applied consistently across document sizes <ref:2606.15020#pg2>. But the authors also have a caution regarding OCR, stating that it’s only effective when the platform commits to visual processing across all document sizes because long documents can trigger fallback to cheaper text extraction routes <ref:2606.15020#pg2>.

Elias: That caveat about OCR highlights the performance trade-off; if you want perfect protection, you need consistent visual processing throughout the entire pipeline, which adds complexity to the system design <ref:2606.15020#pg2>. It also reminds us that these PDF Mirage attacks manipulate font and glyph interpretation to mask content for online services <ref:2606.15020#pg2>.

Priya: So, the implication is that document security in this context requires a layered approach: robust visual processing as a primary defense, combined with mechanisms to monitor ingestion stacks so we can identify where these representation gaps are most likely to occur <ref:2606.15020#pg1>. It’s about securing the entire pipeline upstream of the model reasoning process.

Nadia: That leads us perfectly into the conclusion of this work, "What Users See Is Not What Models Read: Split-View PDFs in Document-to-LLM Supply Chains," and its authors, Side Liu and Jiang Ming <ref:2606.15020#pg0>. They systematically identified twenty-five extraction gaps across four families, proving that document processing layers create divergent semantic views before the model even sees the data <ref:2606.15020#pg1>.

Paper summary: Elias: I think what this paper really contributes is mapping out these specific representation gaps and building that two-tier benchmark to test if those gaps actually cause real downstream issues in commercial LLM services <ref:2606.15020#pg1>. It moves the problem from a vague concern about AI input to a catalog of specific, measurable vulnerabilities in the PDF supply chain <ref:2606.15020#pg1>.

Priya: For privacy research, the implication is that document provenance and integrity checks need to be applied not just at the final model stage but throughout every extraction and rendering step, because hidden text can carry claims or data that are completely invisible to the user <ref:2606.15020#pg1>. It means what we collect from documents isn't just what’s visible on the screen.

Nadia: Exactly, Priya; it really puts the pressure on developers to address these gaps at the parsing and normalization layers rather than just treating them as post-processing issues <ref:2606.15020#pg2>. The practical reality is that if an attacker can control or influence that hidden layer, they can manipulate the AI’s perception of a document without ever changing what a user sees <ref:2606.15020#pg1>.

Elias: And from a cryptographic standpoint, Elias, the assumption here is that the text extracted by the model is not purely what was rendered visually, which means any cryptographic proof relying on visual fidelity of that text might be undermined if it relies on an unverified extraction layer <ref:2606.15020#pg0>. The parameters that break this are those related to font encoding and reading order, as the paper details <ref:2606.15020#pg1>.

Priya: So, when we think about the impact on the world, it’s not just about specific documents being compromised but about a fundamental erosion of trust in how we use AI to interpret complex visual information from documents <ref:2606.15020#pg1>. If the input itself is split into two different realities, the resulting analysis is inherently suspect <ref:2606.15020#pg1>.

Nadia: That’s what we’re seeing; the paper shows that this isn't a niche technical problem, but a widespread exposure across many stacks and services because of these tolerated gaps in the PDF standard itself <ref:2606.15020#pg0>. The challenge moving forward seems to be developing consistency checkers that go beyond simple text size filters to actually verify semantic alignment between the visual output and the machine-consumed text <ref:2606.15020#pg1>.

Conclusion: Nadia: So, to wrap up this discussion on "What Users See Is Not What Models Read: Split-View PDFs in Document-to-LLM Supply Chains," we've seen how document pipelines have these hidden semantic inconsistencies between what a person sees and what the AI actually processes.

Elias: I think that title really gets to the heart of it, focusing on that divergence between rendering and extraction—it suggests a fundamental mismatch in how data flows through this supply chain.

Priya: From my viewpoint, it points toward a significant gap in privacy measurement because if the extracted text differs from the visual text, we can't trust any downstream analysis based on that input.

Nadia: Exactly, Priya; the authors are showing us that this isn't just a rendering hiccup but a systemic failure where attacker-controlled or extractor-dependent text slips through unnoticed.

Elias: That systemic nature is key because it means the vulnerabilities aren't isolated bugs in one piece of software; they’re structural weaknesses permitted by how PDFs are represented.

Priya: And when you consider the scope, this has major implications for document provenance; if a hidden payload can be injected and remain invisible, tracking the true origin of information becomes incredibly difficult.

Nadia: Right, and that brings us to the real-world impact; we're looking at a situation where we might be feeding an AI documents that look clean but have malicious or misleading text embedded in ways only the extractor can see.

Elias: The authors’ work on mapping those twenty-five specific extraction gaps provides a concrete blueprint for understanding exactly which PDF features are most susceptible to this semantic divergence.

Priya: The implications for researchers and developers are that defenses can't just be about blocking visible text; they need checks that verify the consistency between the visual representation and the machine-read data.

Nadia: Precisely, so we’re moving toward building tools that check for semantic alignment across modalities rather than relying on simple heuristics.

Elias: And I'm curious about how those specific font-decoding and reading-order splits you mentioned might affect cryptographic proofs that rely on the assumption of visual fidelity.

Priya: That’s a big question, Elias; if the underlying text representation is mutable based on the extraction method, then any security layer built on top of that text gets fundamentally shaky.

Nadia: It makes us realize that securing this pipeline needs to happen far upstream at the parsing and normalization layers, long before the data ever reaches the reasoning model.

Elias: That structural focus is what makes this paper so valuable; it moves us away from treating these issues as isolated incident reports toward a comprehensive understanding of PDF processing flaws.

More episodes

← Home