From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "From Public Posts to AI-Search Citations".
Nadia: The gist: AI-search citations can turn ordinary web publication into an input path for generated answers,
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So we’re looking at a paper called "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search." It sounds a bit technical, but essentially they are investigating how easily regular web posts can become inputs for answers generated by AI search engines.
Elias: Yeah, that’s the core idea. They are focusing on this selection layer in AI search—how the platform picks and ranks sources before it even shows you anything—and they’re checking if that choice can turn a simple post into something cited in an answer.
Priya: It sounds like they're trying to map out a path from just posting something online to getting your content picked up by an AI search result. That's the big question for privacy and data exposure.
Nadia: Exactly, Priya. They’re saying that if an AI search platform keeps citing domains where it’s really easy for new users to post stuff, then just publishing on those platforms could become a way your content gets pulled into AI-search citations and answer text.
Elias: The authors call this what they call a citation–governance gap, which is basically the difference between where you can actually put content on a website and how easily an AI search platform decides to cite that source.
Priya: That sounds like it could be really problematic because it suggests that control over where you publish doesn't always match control over how an AI system uses that information, which is something we need to figure out for user privacy.
The paper's summary: Nadia: So, what they found in "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search" is that citation patterns aren't random; they get concentrated on a few platforms. They found that about seventy point eight percent of citations are going to just twenty domains on the most popular platform.
Elias: That concentration is key, because when you look at those top sources, they also have low or medium barriers for new users to get an account set up and post content on them.
Priya: So, the paper shows that these frequently cited sources aren't just locked down to big corporations; many of them are platforms where you can jump in with minimal effort. That really makes the idea of a citation path from ordinary posts much more accessible than we thought before.
Nadia: Right, and they tested if content you published on those preferred platforms actually shows up in AI search answers. And they found that ordinary publication on those preferred platforms did change what entered AI search outputs; eight out of ten platforms cited a fabricated concept within seven days in their baseline experiment.
Elias: The measurement itself has some serious challenges, though. They ran into three obstacles: figuring out who is actually responsible for a later citation, dealing with how citations shift over time because pages get indexed at different moments, and knowing which sources to test initially because the platforms don't tell you what they actually cite or how easy it is to publish there.
Priya: I think that attribution challenge is huge. If you publish something and later it gets cited, proving that the citation wasn't just some random index update or another publisher picking it up is going to be really tough for anyone trying to defend their content.
The paper's improvements: Nadia: Now they don’t just stop there; they suggest ways we can actually measure this fragility better. One improvement is that AI search platforms should put in place source-domain diversity enforcement so they have to cite a wider variety of sources instead of just concentrating on a few low-barrier domains.
Elias: That makes sense from a security standpoint. If you force them to cite more diverse sources, you dilute the impact of any single low-barrier platform, which is what we want to stop the concentration effect they found in their initial measurements.
Priya: And then there's the idea of provenance-aware reranking, where content gets downweighted based on its source and how old it is. They suggest this to specifically target that behavior where AI search might give a fast citation right after you post something, which they call "fast citation after publication."
Nadia: That’s smart because it directly addresses the temporal dynamics they found—how things build up and shift across platforms over time. And they also propose developing systems to detect abrupt account-topic shifts or tightly grouped posting times as signs of potential manipulation, like Generative Engine Optimization activity.
Elias: And then there's the UI part, which is integrating transparency signals into the user interface to flag citations that look unusually concentrated or dominated by platforms that are known for being easy to publish on. That gives users context about where their answer might be coming from.
Priya: I think those improvements focus a lot on building better detection tools rather than just understanding the problem itself, which is important because the underlying structural finding is that this citation–governance gap exists in the first place.
Conclusion: Nadia: So to wrap up on "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search," they are saying that citation security isn't just about one single piece of defense; it’s a whole supply chain problem involving checking the source, checking the path you took to publish, and looking at how long it takes for that citation to appear.
Elias: They conclude that AI search citations can be an input path for generated answers directly from ordinary web publication, which means we need defenses focused on auditing where those pages come from and what the entire publication path looks like.
Priya: For me, this really highlights the tension between how content is created and how it’s used by these massive AI systems; it’s about making sure the control you have over your own publishing activity actually translates into safety when an AI system starts using that content as a foundation for its answers.
Nadia: Exactly. We need to look at the whole chain, from the moment you post something online to when an answer is synthesized by AI search platforms. That’s where the security risk lies in this paper's findings on "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search."
Elias: It leaves us with a clear direction for defense—we need to focus on auditing that source provenance and tracking those citation latencies. That’s our path forward.
Priya: I think it’s a lot to digest, but understanding this gap between publication control and AI search visibility is crucial for anyone building systems on top of these search engines.
Qi Liu†, Geng Hong†✉, Xinyang Zhang†, Pei Chen†, Yutong Li†, Min Yang†‡✉
Fudan University, China
cs.CR
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/qiqiqi-roxie/query-list
Project page: https://zhituishidai.com
License: http://creativecommons.org/licenses/by/4.0/
The gist: The gist: AI-search citations can turn ordinary web publication into an input path for generated answers, creating a security problem where source choice becomes a security question How it works
Key concepts
- Retrieval-Augmented Generation (RAG)
- This is the system AI search uses to generate answers. It works by first retrieving relevant web pages, then filtering and ranking them, and finally synthesizing an answer with citations attached. This process makes the choice of which source to cite a security concern.
- Citation-Governance Gap
- This gap exists when a domain is heavily cited by AI search but the platform controlling that domain has low barriers for new publishers to post content. This mismatch means ordinary web publishing can create AI search visibility without needing access to the original source's controlled platform.
- Publication Barrier Testing
- This involved testing how easy it is for new users to set up accounts and post content on platforms associated with frequently cited sources. The study found that many such platforms have low or medium barriers, meaning high-volume sources are not always limited to tightly controlled publishers.
Terminology
Summary
The gist: AI-search citations can turn ordinary web publication into an input path for generated answers, creating a security problem where source choice becomes a security question
How it works
AI-search platforms function using a Retrieval-Augmented Generation (RAG) pipeline where they retrieve candidate pages, filter and rank sources, and present a synthesized answer with citations attached This selection layer may amplify source bias and turn source choice into a security question The authors define a citation–governance gap as a mismatch between citation visibility and publication control, where an AI-search platform repeatedly cites a source domain while the ability to place content on that domain is governed by a separate publication platform with low barriers for new publishers This gap creates a practical path from ordinary web publication to AI-search citation, requiring only ordinary publication access to public web platforms
Citation Concentration and Publication Access
The research first measures where AI-search citations come from, finding that citation distributions are concentrated Across 10 platforms, 70.8% of citations go to just 20 domains on the most concentrated platform The study identifies a citation–governance gap when a frequently cited domain also has low publication barriers Publication barrier testing revealed that 15 of 22 tested publication platforms tied to cited source domains had low or medium barriers for both account setup and posting Finding 2 shows that high-volume and broadly cited sources are not limited to tightly controlled publishers, as many map to platforms that new accounts can enter with low or medium setup barriers
Targeted Publication in AI Search
The second research question tests whether newly published content on selected publication platforms becomes visible in AI-search citations and answer text The authors track two observable signals: explicit citation of a published URL and appearance of a publication-specific marker in answer text In the baseline experiment, ordinary publication on RQ1-preferred platforms changed what entered AIsearch outputs, with 8 of 10 platforms citing a fabricated concept within seven days Finding 3 shows that under our measurement conditions, ordinary publishers gained AI-search visibility by publishing on citation-preferred platforms without AI-search access
Temporal Dynamics and GEO Services
The third research question examines the temporal dynamics of targeted publication, showing that citation visibility is not a onetime event but builds, shifts across platforms, and may remain after deletion The study also tests whether this publication-to-citation path is already used in the wild through Generative Engine Optimization (GEO) services A single GEO purchase produced one confirmed commercial workflow linking payment, public posting, and production AI-search citation with embedded markers
Conclusion
Citation security is therefore a publication-supply-chain problem: defenses should audit citedpage provenance, publication path, and citation latency The paper concludes that AI-search citations can turn ordinary web publication into an input path for generated answers
ETHICS CONSIDERATIONS
Conducting targeted-publication experiments on production AI-search platforms requires careful attention to potential harm and platform policy constraints The authors designed the study with reference to the Menlo Report’s principles for ICT research and implemented safeguards throughout the measurement pipeline The authors used generative AI tools in three ways, including using ChatGPT to seed candidate query wording for the RQ1 query panel All experimental topics were chosen under a strict harmavoidance policy, excluding high-stakes domains such as health and medical advice or political content The authors completed responsible disclosure to all 10 AI-search platforms describing their methodology and key findings The authors acknowledge that the quantitative results reported here are likely to be more valuable to defenders than to would-be manipulators The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat the results as structural findings rather than claims that are likely to drift with platform updates The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on
Improvements for AI systems
-
AI-search platforms should implement source-domain diversity enforcement to limit exposure, as
Diversity can limit any single domain’s contribution but may hurt narrow queries that legitimately rely on specialized sources
(Section VII.B). This prevents concentration of AI-search citations on a few low-barrier domains by requiring the system to cite a broader set of sources. -
Implement provenance-aware reranking to downweight content based on its publication source and temporal signals, as
Provenance and temporal signals can downweight recently published UGC or detect same-domain clustering, synchronized publication, and fast citation after publication
(Section VII.B). This targets thefast citation after publication
behavior observed in the n=1 case study. -
Develop a system to detect
abrupt account-topic shifts, tightly grouped posting times, near-duplicate content,
as these signals are identified as indicators of potential GEO activity (Section VII.B). This helps in identifyingabrupt account-topic shifts
which are noted as a signal for prioritizing citation channels for review. -
Integrate transparency signals into the user interface to flag citations that are
unusually concentrated, unusually recent, or dominated by low-barrier publication platforms
(Section VII.B). This allows end users to be aware of the specific source of their synthesized answer, improvingcitation context awareness.
-
For content generation pipelines, incorporate a mechanism to distinguish between explicit citation and marker appearance in generated answers (Section V-B). This ensures that the system distinguishes between
explicit citation of a published URL and appearance of a publication-specific marker in answer text,
which are distinct signals.
Abstract
As more users ask AI systems for information, AI-search platforms are becoming a common gateway to web information. Unlike traditional search, which maps keywords to ranked pages, AI search retrieves pages, filters sources, selects citations, and generates answers before users see sources. This selection layer may amplify source bias and turn source choice into a security question. If a platform repeatedly cites domains where new users can publish posts easily, ordinary publication on those domains can become an indirect path into AI-search citations and answer text. Measuring this path is hard: platforms reveal little about citation selection, citations change over time, and the web contains so much background content that later answer changes are hard to attribute to our posts. We present a measurement framework for identifying and measuring this low-barrier publication path, combining cross-platform citation mapping, publication-barrier testing, and marker-controlled publication experiments. Across 10 AI-search platforms, we analyze 17,211 citation instances over 6,356 unique source domains and find: (1) citations concentrate in platform-specific sources, with top-20 domains capturing 20.5--70.8% of per-platform citations, and 15 of 22 tested publication platforms tied to cited source domains had low or medium barriers for both account setup and posting; (2) in our experiments, ordinary publication on preferred platforms changed what entered AI-search outputs: 8 of 10 platforms cited a fabricated concept within seven days, and one high-preference-platform article had greater citation impact than over 20 matched low-preference posts; and (3) this path is commercially available: a 14 GEO purchase produced 13 public posts, and one AI-search platform cited GEO-posted content with our designed markers within one hour.
Sources
- The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale
- Search Engines Post-ChatGPT: How Generative Artificial Intelligence Could Make Search Less Reliable
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact
- From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms
- Answer Bubbles: Information Exposure in AI-Mediated Search
- Adversarial Search Engine Optimization for Large Language Models
- Dynamics of Adversarial Attacks on Large Language Model-Based Search Engines
- BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
- TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
- Ignore Previous Prompt: Attack Techniques For Language Models
- Certifiably Robust RAG against Retrieval Corruption
- Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs