From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search
summary
The gist
The gist: AI-search citations can turn ordinary web publication into an input path for generated answers, creating a security problem where source choice becomes a security question How it works
In short
The research investigated how ordinary web publications can become inputs for AI search answers through a Retrieval-Augmented Generation (RAG) pipeline. It found a 'citation-governance gap' where frequently cited domains have low publication barriers, allowing new publishers to gain visibility in AI search results simply by posting content on those platforms.
Key concepts
- Retrieval-Augmented Generation (RAG)
- This is the system AI search uses to generate answers. It works by first retrieving relevant web pages, then filtering and ranking them, and finally synthesizing an answer with citations attached. This process makes the choice of which source to cite a security concern.
- Citation-Governance Gap
- This gap exists when a domain is heavily cited by AI search but the platform controlling that domain has low barriers for new publishers to post content. This mismatch means ordinary web publishing can create AI search visibility without needing access to the original source's controlled platform.
- Publication Barrier Testing
- This involved testing how easy it is for new users to set up accounts and post content on platforms associated with frequently cited sources. The study found that many such platforms have low or medium barriers, meaning high-volume sources are not always limited to tightly controlled publishers.
Terminology used across episodes
This episode discusses
- From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search · Paper Radio
- The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale
- Search Engines Post-ChatGPT: How Generative Artificial Intelligence Could Make Search Less Reliable
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact
- From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms
- Answer Bubbles: Information Exposure in AI-Mediated Search · Paper Radio
- Adversarial Search Engine Optimization for Large Language Models
- Dynamics of Adversarial Attacks on Large Language Model-Based Search Engines
- BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models · Paper Radio
- TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
- Ignore Previous Prompt: Attack Techniques For Language Models
- Certifiably Robust RAG against Retrieval Corruption
- Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
The paper
From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search · Read on arXiv
Qi Liu†, Geng Hong†✉, Xinyang Zhang†, Pei Chen†, Yutong Li†, Min Yang†‡✉
Fudan University, China
As more users ask AI systems for information, AI-search platforms are becoming a common gateway to web information. Unlike traditional search, which maps keywords to ranked pages, AI search retrieves pages, filters sources, selects citations, and generates answers before users see sources. This selection layer may amplify source bias and turn source choice into a security question. If a platform repeatedly cites domains where new users can publish posts easily, ordinary publication on those domains can become an indirect path into AI-search citations and answer text. Measuring this path is hard: platforms reveal little about citation selection, citations change over time, and the web contains so much background content that later answer changes are hard to attribute to our posts. We present a measurement framework for identifying and measuring this low-barrier publication path, combining cross-platform citation mapping, publication-barrier testing, and marker-controlled publication experiments. Across 10 AI-search platforms, we analyze 17,211 citation instances over 6,356 unique source domains and find: (1) citations concentrate in platform-specific sources, with top-20 domains capturing 20.5--70.8% of per-platform citations, and 15 of 22 tested publication platforms tied to cited source domains had low or medium barriers for both account setup and posting; (2) in our experiments, ordinary publication on preferred platforms changed what entered AI-search outputs: 8 of 10 platforms cited a fabricated concept within seven days, and one high-preference-platform article had greater citation impact than over 20 matched low-preference posts; and (3) this path is commercially available: a 14 GEO purchase produced 13 public posts, and one AI-search platform cited GEO-posted content with our designed markers within one hour.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "From Public Posts to AI-Search Citations".
Nadia: The gist: AI-search citations can turn ordinary web publication into an input path for generated answers,
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So we’re looking at a paper called "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search." It sounds a bit technical, but essentially they are investigating how easily regular web posts can become inputs for answers generated by AI search engines.
Elias: Yeah, that’s the core idea. They are focusing on this selection layer in AI search—how the platform picks and ranks sources before it even shows you anything—and they’re checking if that choice can turn a simple post into something cited in an answer.
Priya: It sounds like they're trying to map out a path from just posting something online to getting your content picked up by an AI search result. That's the big question for privacy and data exposure.
Nadia: Exactly, Priya. They’re saying that if an AI search platform keeps citing domains where it’s really easy for new users to post stuff, then just publishing on those platforms could become a way your content gets pulled into AI-search citations and answer text.
Elias: The authors call this what they call a citation–governance gap, which is basically the difference between where you can actually put content on a website and how easily an AI search platform decides to cite that source.
Priya: That sounds like it could be really problematic because it suggests that control over where you publish doesn't always match control over how an AI system uses that information, which is something we need to figure out for user privacy.
The paper's summary: Nadia: So, what they found in "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search" is that citation patterns aren't random; they get concentrated on a few platforms. They found that about seventy point eight percent of citations are going to just twenty domains on the most popular platform.
Elias: That concentration is key, because when you look at those top sources, they also have low or medium barriers for new users to get an account set up and post content on them.
Priya: So, the paper shows that these frequently cited sources aren't just locked down to big corporations; many of them are platforms where you can jump in with minimal effort. That really makes the idea of a citation path from ordinary posts much more accessible than we thought before.
Nadia: Right, and they tested if content you published on those preferred platforms actually shows up in AI search answers. And they found that ordinary publication on those preferred platforms did change what entered AI search outputs; eight out of ten platforms cited a fabricated concept within seven days in their baseline experiment.
Elias: The measurement itself has some serious challenges, though. They ran into three obstacles: figuring out who is actually responsible for a later citation, dealing with how citations shift over time because pages get indexed at different moments, and knowing which sources to test initially because the platforms don't tell you what they actually cite or how easy it is to publish there.
Priya: I think that attribution challenge is huge. If you publish something and later it gets cited, proving that the citation wasn't just some random index update or another publisher picking it up is going to be really tough for anyone trying to defend their content.
The paper's improvements: Nadia: Now they don’t just stop there; they suggest ways we can actually measure this fragility better. One improvement is that AI search platforms should put in place source-domain diversity enforcement so they have to cite a wider variety of sources instead of just concentrating on a few low-barrier domains.
Elias: That makes sense from a security standpoint. If you force them to cite more diverse sources, you dilute the impact of any single low-barrier platform, which is what we want to stop the concentration effect they found in their initial measurements.
Priya: And then there's the idea of provenance-aware reranking, where content gets downweighted based on its source and how old it is. They suggest this to specifically target that behavior where AI search might give a fast citation right after you post something, which they call "fast citation after publication."
Nadia: That’s smart because it directly addresses the temporal dynamics they found—how things build up and shift across platforms over time. And they also propose developing systems to detect abrupt account-topic shifts or tightly grouped posting times as signs of potential manipulation, like Generative Engine Optimization activity.
Elias: And then there's the UI part, which is integrating transparency signals into the user interface to flag citations that look unusually concentrated or dominated by platforms that are known for being easy to publish on. That gives users context about where their answer might be coming from.
Priya: I think those improvements focus a lot on building better detection tools rather than just understanding the problem itself, which is important because the underlying structural finding is that this citation–governance gap exists in the first place.
Conclusion: Nadia: So to wrap up on "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search," they are saying that citation security isn't just about one single piece of defense; it’s a whole supply chain problem involving checking the source, checking the path you took to publish, and looking at how long it takes for that citation to appear.
Elias: They conclude that AI search citations can be an input path for generated answers directly from ordinary web publication, which means we need defenses focused on auditing where those pages come from and what the entire publication path looks like.
Priya: For me, this really highlights the tension between how content is created and how it’s used by these massive AI systems; it’s about making sure the control you have over your own publishing activity actually translates into safety when an AI system starts using that content as a foundation for its answers.
Nadia: Exactly. We need to look at the whole chain, from the moment you post something online to when an answer is synthesized by AI search platforms. That’s where the security risk lies in this paper's findings on "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search."
Elias: It leaves us with a clear direction for defense—we need to focus on auditing that source provenance and tracking those citation latencies. That’s our path forward.
Priya: I think it’s a lot to digest, but understanding this gap between publication control and AI search visibility is crucial for anyone building systems on top of these search engines.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits