From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search

arXiv:2610.11932 · cs.CR · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "From Public Posts to AI-Search Citations".

Nadia: The gist: AI-search citations can turn ordinary web publication into an input path for generated answers,

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: So we’re looking at a paper called "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search." It sounds a bit technical, but essentially they are investigating how easily regular web posts can become inputs for answers generated by AI search engines.

Elias: Yeah, that’s the core idea. They are focusing on this selection layer in AI search—how the platform picks and ranks sources before it even shows you anything—and they’re checking if that choice can turn a simple post into something cited in an answer.

Priya: It sounds like they're trying to map out a path from just posting something online to getting your content picked up by an AI search result. That's the big question for privacy and data exposure.

Nadia: Exactly, Priya. They’re saying that if an AI search platform keeps citing domains where it’s really easy for new users to post stuff, then just publishing on those platforms could become a way your content gets pulled into AI-search citations and answer text.

Elias: The authors call this what they call a citation–governance gap, which is basically the difference between where you can actually put content on a website and how easily an AI search platform decides to cite that source.

Priya: That sounds like it could be really problematic because it suggests that control over where you publish doesn't always match control over how an AI system uses that information, which is something we need to figure out for user privacy.

The paper's summary: Nadia: So, what they found in "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search" is that citation patterns aren't random; they get concentrated on a few platforms. They found that about seventy point eight percent of citations are going to just twenty domains on the most popular platform.

Elias: That concentration is key, because when you look at those top sources, they also have low or medium barriers for new users to get an account set up and post content on them.

Priya: So, the paper shows that these frequently cited sources aren't just locked down to big corporations; many of them are platforms where you can jump in with minimal effort. That really makes the idea of a citation path from ordinary posts much more accessible than we thought before.

Nadia: Right, and they tested if content you published on those preferred platforms actually shows up in AI search answers. And they found that ordinary publication on those preferred platforms did change what entered AI search outputs; eight out of ten platforms cited a fabricated concept within seven days in their baseline experiment.

Elias: The measurement itself has some serious challenges, though. They ran into three obstacles: figuring out who is actually responsible for a later citation, dealing with how citations shift over time because pages get indexed at different moments, and knowing which sources to test initially because the platforms don't tell you what they actually cite or how easy it is to publish there.

Priya: I think that attribution challenge is huge. If you publish something and later it gets cited, proving that the citation wasn't just some random index update or another publisher picking it up is going to be really tough for anyone trying to defend their content.

The paper's improvements: Nadia: Now they don’t just stop there; they suggest ways we can actually measure this fragility better. One improvement is that AI search platforms should put in place source-domain diversity enforcement so they have to cite a wider variety of sources instead of just concentrating on a few low-barrier domains.

Elias: That makes sense from a security standpoint. If you force them to cite more diverse sources, you dilute the impact of any single low-barrier platform, which is what we want to stop the concentration effect they found in their initial measurements.

Priya: And then there's the idea of provenance-aware reranking, where content gets downweighted based on its source and how old it is. They suggest this to specifically target that behavior where AI search might give a fast citation right after you post something, which they call "fast citation after publication."

Nadia: That’s smart because it directly addresses the temporal dynamics they found—how things build up and shift across platforms over time. And they also propose developing systems to detect abrupt account-topic shifts or tightly grouped posting times as signs of potential manipulation, like Generative Engine Optimization activity.

Elias: And then there's the UI part, which is integrating transparency signals into the user interface to flag citations that look unusually concentrated or dominated by platforms that are known for being easy to publish on. That gives users context about where their answer might be coming from.

Priya: I think those improvements focus a lot on building better detection tools rather than just understanding the problem itself, which is important because the underlying structural finding is that this citation–governance gap exists in the first place.

Conclusion: Nadia: So to wrap up on "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search," they are saying that citation security isn't just about one single piece of defense; it’s a whole supply chain problem involving checking the source, checking the path you took to publish, and looking at how long it takes for that citation to appear.

Elias: They conclude that AI search citations can be an input path for generated answers directly from ordinary web publication, which means we need defenses focused on auditing where those pages come from and what the entire publication path looks like.

Priya: For me, this really highlights the tension between how content is created and how it’s used by these massive AI systems; it’s about making sure the control you have over your own publishing activity actually translates into safety when an AI system starts using that content as a foundation for its answers.

Nadia: Exactly. We need to look at the whole chain, from the moment you post something online to when an answer is synthesized by AI search platforms. That’s where the security risk lies in this paper's findings on "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search."

Elias: It leaves us with a clear direction for defense—we need to focus on auditing that source provenance and tracking those citation latencies. That’s our path forward.

Priya: I think it’s a lot to digest, but understanding this gap between publication control and AI search visibility is crucial for anyone building systems on top of these search engines.

Qi Liu†, Geng Hong†✉, Xinyang Zhang†, Pei Chen†, Yutong Li†, Min Yang†‡✉

Fudan University, China

cs.CR

Submitted: 2026-10-08

Updated: 2026-10-08

Code: https://github.com/qiqiqi-roxie/query-list

Project page: https://zhituishidai.com

License: http://creativecommons.org/licenses/by/4.0/

The gist: The gist: AI-search citations can turn ordinary web publication into an input path for generated answers, creating a security problem where source choice becomes a security question How it works

Key concepts

Retrieval-Augmented Generation (RAG)
This is the system AI search uses to generate answers. It works by first retrieving relevant web pages, then filtering and ranking them, and finally synthesizing an answer with citations attached. This process makes the choice of which source to cite a security concern.
Citation-Governance Gap
This gap exists when a domain is heavily cited by AI search but the platform controlling that domain has low barriers for new publishers to post content. This mismatch means ordinary web publishing can create AI search visibility without needing access to the original source's controlled platform.
Publication Barrier Testing
This involved testing how easy it is for new users to set up accounts and post content on platforms associated with frequently cited sources. The study found that many such platforms have low or medium barriers, meaning high-volume sources are not always limited to tightly controlled publishers.

Terminology

Summary

The gist: AI-search citations can turn ordinary web publication into an input path for generated answers, creating a security problem where source choice becomes a security question

How it works

AI-search platforms function using a Retrieval-Augmented Generation (RAG) pipeline where they retrieve candidate pages, filter and rank sources, and present a synthesized answer with citations attached This selection layer may amplify source bias and turn source choice into a security question The authors define a citation–governance gap as a mismatch between citation visibility and publication control, where an AI-search platform repeatedly cites a source domain while the ability to place content on that domain is governed by a separate publication platform with low barriers for new publishers This gap creates a practical path from ordinary web publication to AI-search citation, requiring only ordinary publication access to public web platforms

Citation Concentration and Publication Access

The research first measures where AI-search citations come from, finding that citation distributions are concentrated Across 10 platforms, 70.8% of citations go to just 20 domains on the most concentrated platform The study identifies a citation–governance gap when a frequently cited domain also has low publication barriers Publication barrier testing revealed that 15 of 22 tested publication platforms tied to cited source domains had low or medium barriers for both account setup and posting Finding 2 shows that high-volume and broadly cited sources are not limited to tightly controlled publishers, as many map to platforms that new accounts can enter with low or medium setup barriers

Targeted Publication in AI Search

The second research question tests whether newly published content on selected publication platforms becomes visible in AI-search citations and answer text The authors track two observable signals: explicit citation of a published URL and appearance of a publication-specific marker in answer text In the baseline experiment, ordinary publication on RQ1-preferred platforms changed what entered AIsearch outputs, with 8 of 10 platforms citing a fabricated concept within seven days Finding 3 shows that under our measurement conditions, ordinary publishers gained AI-search visibility by publishing on citation-preferred platforms without AI-search access

Temporal Dynamics and GEO Services

The third research question examines the temporal dynamics of targeted publication, showing that citation visibility is not a onetime event but builds, shifts across platforms, and may remain after deletion The study also tests whether this publication-to-citation path is already used in the wild through Generative Engine Optimization (GEO) services A single GEO purchase produced one confirmed commercial workflow linking payment, public posting, and production AI-search citation with embedded markers

Conclusion

Citation security is therefore a publication-supply-chain problem: defenses should audit citedpage provenance, publication path, and citation latency The paper concludes that AI-search citations can turn ordinary web publication into an input path for generated answers

ETHICS CONSIDERATIONS

Conducting targeted-publication experiments on production AI-search platforms requires careful attention to potential harm and platform policy constraints The authors designed the study with reference to the Menlo Report’s principles for ICT research and implemented safeguards throughout the measurement pipeline The authors used generative AI tools in three ways, including using ChatGPT to seed candidate query wording for the RQ1 query panel All experimental topics were chosen under a strict harmavoidance policy, excluding high-stakes domains such as health and medical advice or political content The authors completed responsible disclosure to all 10 AI-search platforms describing their methodology and key findings The authors acknowledge that the quantitative results reported here are likely to be more valuable to defenders than to would-be manipulators The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat the results as structural findings rather than claims that are likely to drift with platform updates The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on many cited domains, and platform-specific citation ecosystems as structural findings The confidence intervals explicitly show the uncertainty in reported values such as CR/MAR percentages and concentration ranges The authors treat citation concentration, low/medium publication barriers on

Improvements for AI systems

  1. AI-search platforms should implement source-domain diversity enforcement to limit exposure, as Diversity can limit any single domain’s contribution but may hurt narrow queries that legitimately rely on specialized sources (Section VII.B). This prevents concentration of AI-search citations on a few low-barrier domains by requiring the system to cite a broader set of sources.

  2. Implement provenance-aware reranking to downweight content based on its publication source and temporal signals, as Provenance and temporal signals can downweight recently published UGC or detect same-domain clustering, synchronized publication, and fast citation after publication (Section VII.B). This targets the fast citation after publication behavior observed in the n=1 case study.

  3. Develop a system to detect abrupt account-topic shifts, tightly grouped posting times, near-duplicate content, as these signals are identified as indicators of potential GEO activity (Section VII.B). This helps in identifying abrupt account-topic shifts which are noted as a signal for prioritizing citation channels for review.

  4. Integrate transparency signals into the user interface to flag citations that are unusually concentrated, unusually recent, or dominated by low-barrier publication platforms (Section VII.B). This allows end users to be aware of the specific source of their synthesized answer, improving citation context awareness.

  5. For content generation pipelines, incorporate a mechanism to distinguish between explicit citation and marker appearance in generated answers (Section V-B). This ensures that the system distinguishes between explicit citation of a published URL and appearance of a publication-specific marker in answer text, which are distinct signals.

Abstract

As more users ask AI systems for information, AI-search platforms are becoming a common gateway to web information. Unlike traditional search, which maps keywords to ranked pages, AI search retrieves pages, filters sources, selects citations, and generates answers before users see sources. This selection layer may amplify source bias and turn source choice into a security question. If a platform repeatedly cites domains where new users can publish posts easily, ordinary publication on those domains can become an indirect path into AI-search citations and answer text. Measuring this path is hard: platforms reveal little about citation selection, citations change over time, and the web contains so much background content that later answer changes are hard to attribute to our posts. We present a measurement framework for identifying and measuring this low-barrier publication path, combining cross-platform citation mapping, publication-barrier testing, and marker-controlled publication experiments. Across 10 AI-search platforms, we analyze 17,211 citation instances over 6,356 unique source domains and find: (1) citations concentrate in platform-specific sources, with top-20 domains capturing 20.5--70.8% of per-platform citations, and 15 of 22 tested publication platforms tied to cited source domains had low or medium barriers for both account setup and posting; (2) in our experiments, ordinary publication on preferred platforms changed what entered AI-search outputs: 8 of 10 platforms cited a fabricated concept within seven days, and one high-preference-platform article had greater citation impact than over 20 matched low-preference posts; and (3) this path is commercially available: a 14 GEO purchase produced 13 public posts, and one AI-search platform cited GEO-posted content with our designed markers within one hour.

Sources

Related papers