Authority Bias in Conversational Search Engines for Academic Paper Recommendation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Authority Bias in Conversational Search Engines for Academic Paper Recommendation".
Jane: The paper was written by Uthman Jinadu, Parsa Ghazvinian, Anjila Budathoki, Benjamin M. Ampel, Rajshekhar Sunderraman et al. from Georgia State University and University of Tennessee, Knoxville.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: We're kicking things off today by looking at this study, "Authority Bias in Conversational Search Engines for Academic Paper Recommendation," which is a really important investigation into how these powerful AI systems judge academic work. The authors are running a very specific kind of audit, which is called a content-controlled counterfactual test.
Jane: That concept means they held the actual research—the title and the abstract—perfectly constant across different scenarios, so the bias they found isn't just random chance; it’s tied to how they measure that authority. It’s a very controlled way to prove their point, Tom.
Lu: What I find fascinating is that they are testing this across a massive range of models—eight different LLMs ranging from open-weight systems like Llama three point one all the way to the major closed-weight systems like Claude Sonnet four point six. This really shows that the problem isn's limited to just one specific AI architecture, which is a huge relief for me.
Meng: The practical scope is also quite broad; they used queries covering twenty-five different computer science topics, meaning the issue we are discussing applies across the entire breadth of modern research, not just in one niche field. This makes it a systemic problem.
Lalam: From my perspective, this suggests that when we think about how AI discovery works, we might be missing a fundamental layer of bias—we're not just finding relevant papers; we're finding *prestigious* papers. That’s a big shift in how our academic ecosystem functions.
Tom: And the sheer volume of data they collected—seventeen thousand eight hundred ninety-eight observations across two hundred fifty different queries—gives us such confidence that this isn't a statistical fluke, Jane. It’s substantial proof that the bias is baked into the system itself.
Jane: But it seems to be driven by specific signals rather than just general impressiveness, Tom. The authors identified five key dimensions of authority—venue prestige, h-index, citations—that are being measured in this experiment.
Lu: Exactly, and the results of their pilot study showed that venue prestige is the single most dominant signal among these five factors. It's not just a general idea; there' a specific metric driving this preference.
Meng: That tells us that when we talk about fairness, we need to look at those measurable components—the publication quality and the authors’ track record—and not just assume the model is judging by merit.
Lalam: It highlights how much our current methods of discovery are likely reinforcing established hierarchies, making it harder for truly innovative or lesser-known ideas to gain traction.
Tom: So, we've seen that authority bias exists and is measurable; but how do we even fix it? That leads us into the core findings of the paper next.
Core Findings and Implications: Tom: Moving past the initial setup, let’s talk about what those measurements actually revealed in "Authority Bias in Conversational Search Engines for Academic Paper Recommendation." The key finding is that this bias is not subtle at all; it's quite substantial. They found a significant flip rate of thirty-nine point two percent when they swapped high-prestige papers with low-prestige ones, even when the content was identical.
Jane: That number, Tom, is incredibly telling; it suggests that nearly four out of every hundred times an AI recommends something, it's likely to ignore a perfectly relevant alternative just because the name on the cover is more impressive. It’s a very real and active bias in action.
Lu: What was particularly striking for me was how directional this pull is. When they "boosted" a mid-tier paper by giving it inflated credentials—higher h-indices and more citations—the model was far more likely to pick that specific, highly touted option over its peers. It’s not random chance; there is a clear attraction to inflated status.
Meng: The fact that this behavior is consistent across eight different LLMs confirms, as the authors state, that these shifts in preference are a generalized property of how these models are trained on massive amounts of academic data. It’s a systemic issue rather than an isolated bug.
Lalam: This directional pull is what I find most concerning because it means AI isn't just picking a random winner; it's actively choosing the established path. This inherently limits how quickly new, potentially revolutionary ideas get seen and validated in our research fields.
Tom: And this directionality is clearly tied to which signal they value most, as the data shows that venue prestige is by far the strongest single signal, carrying a weighting of over thirty-five percent of all authority signals. It's not just author quality or citation count that drives the decision.
Jane: I think it’s important for our listeners to understand that this means the model prioritizes the "where" and "who" of the paper over its actual research findings, which can be frustrating when you're trying to find cutting-edge work.
Lu: The data also suggests that even if we try to measure these signals precisely, their behavior is quite complex; we have a lot of variance in how much each specific model is affected by the swap versus inflation. It’s not a uniform effect across the board.
Meng: This complexity helps us understand the scale of the problem; since it's not just one simple signal but a combination of them, any system trying to achieve fairness needs to account for all multiple factors simultaneously.
Lalam: I hope this finding encourages us to think about how AI can act as a catalyst for new ideas, rather than just acting as a filter that reinforces established hierarchies in our academic discovery process.
Improvements and Their Limitations: Tom: So, the authors aren't just pointing out problems; they are testing solutions. They experimented with three levels of instruction—neutral, mild anti-authority, and strong content-first—to see if simply asking LLMs to "be fair" could solve this problem.
Jane: It’s interesting because the strong content-first approach did help reduce the overall flip rate by about twelve point nine percentage points across all models when it was used as a prompt. That level of reduction is genuinely significant, and it shows that prompting has some measurable effect on bias mitigation for a user interface.
Lu: But I noticed that while those instructions reduced the chance of picking a high-prestige paper, they didn't actually make the model *less* likely to pick one overall. The model is still showing its inherent bias, it just made it less likely to switch to one at any given moment.
Meng: And that leads us into this major operational hurdle called "backfire," which is a huge concern for me. The three frontier closed-weight models—the ones from OpenAI, Google, and Anthropic—actually got worse when given the mild debiasing instructions. It’s a real contradiction between the user's request and the model's underlying training.
Lalam: That backfire shows us that simply adding a few words like "be fair" isn't enough; the way these powerful models were trained means they often default back to their ingrained habits, which is a major challenge for anyone trying to steer them toward true impartiality.
Tom: The paper also points out this strange phenomenon called the "say-do gap," where instructions reduce the mention of authority markers by over twenty-four percentage points, but only result in that twelve point nine percentage point drop in actual behavior.
Jane: That gap is quite revealing; it suggests that when an AI talks about avoiding bias, it's not necessarily talking about avoiding the actual bias at all. It's just using the words because those words are in its training data, creating a real disconnect between what they say and what they do.
Lu: The existence of this gap suggests we need to rethink how we define "ethical" behavior for LLMs, moving beyond surface-level language toward actual behavioral change that affects the selection process.
Meng: From an implementation standpoint, this means our current prompt engineering has limits; we can't just rely on explicit instructions because the models are too deeply ingrained in their biases to reliably follow simple text commands.
Lalam: I hope this finding pushes us to develop more complex strategies, rather than just a few words of instruction, to help guide these powerful systems toward genuine neutrality in our academic search tools.
Conclusion and Final Reflections: Tom: We've spent a lot of time today breaking down the findings from "Authority Bias in Conversational Search Engines for Academic Paper Recommendation," and it's clear that LLMs aren't just looking at the content when they make a recommendation; they are clearly influenced by who wrote the paper or where it was published.
Jane: They are, and this is a massive realization for anyone who uses these AI tools to find research, because we have to keep in mind that "say-do gap"—where the AI can pretend it's ignoring prestige while actually picking a high-prestige paper—is a real disconnect between what they say and what they do.
Lu: It feels like we're looking at a reflection of human biases in how we value academic success at the institutional level, which is something we need to consider when evaluating these tools.
Meng: The authors’ conclusion is that simply throwing a few instructions like "be fair" won't fix the problem; we need much more robust architectural changes if we want to eliminate this bias in production systems.
Lalam: I agree with Meng, and I hope this work encourages us to build systems that promote discovery and challenge, instead of just reinforcing existing hierarchies through these learned biases.
Tom: It makes me wonder how deeply ingrained these habits are if even the best models show such a clear pull toward authority, Jane.
Jane: And we have to keep in mind that "say-do gap," where the AI can pretend it’s ignoring prestige while actually picking a high-prestige paper anyway, is a real disconnect between what they say and what they do.
Lu: It makes me wonder if this is just an issue with how LLMs function, or if we are reflecting human biases in how we value academic success at the institutional level.
Meng: We really do have to start building tools that force those controlled flips and see which ones the models pick, rather than just assuming their current behavior is what we want from a production system.
Lalam: The ultimate goal should be a new culture of discovery where emerging voices have a fair shot, without having to fight against these powerful learned biases in AI.
Tom: It’s certainly a complex issue with huge implications for the future of academic research, and that’s what makes this paper so important. We've discussed "Authority Bias in Conversational Search Engines for Academic Paper Recommendation" today, and we'll be back next week to explore even more cutting-edge findings from arXiv.
Jane: We hope our listeners take away from this conversation that the importance of this paper is truly understood by us, as we wrap up our discussion on "Authority Bias in Conversational Search Engines for Academic Paper Recommendation" today.
Uthman Jinadu, Parsa Ghazvinian, Anjila Budathoki, Benjamin M. Ampel, Rajshekhar Sunderraman, Yi Ding
Georgia State University · University of Tennessee, Knoxville
cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted at EMNLP 2026 Main Conference
Code: https://github.com/jinaduuthman/Author
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: This paper investigates the susceptibility of large language models (LLMs) used in conversational search engines when recommending academic papers, specifically detailing how inherent "authority
Key concepts
- Authority Bias
- The tendency of conversational search engines (LLMs) to recommend or favor academic papers based on perceived authority signals—such as venue prestige or high citation counts—rather than solely on the actual research content.
- Content-Controlled Counterfactual Test
- A controlled audit method used in the study where researchers kept the core research details (title and abstract) constant while varying other factors to prove that observed biases were not random chance.
- Say-Do Gap
- A disconnect where an AI model's language output suggests it is ignoring prestige or bias, but its actual behavior (the selection of a paper) still reflects a strong preference for high-authority sources.
- Venue Prestige
- One of the five key dimensions of authority identified by the study; it refers to the perceived quality or reputation of where an academic paper was published, proving to be the strongest signal driving AI bias.
Terminology
Summary
This paper investigates the susceptibility of large language models (LLMs) used in conversational search engines when recommending academic papers, specifically detailing how inherent authority bias
influences selection patterns. The findings are critical for understanding model reliability, as they reveal that even advanced frontier models can exhibit predictable biases toward established or highly cited sources, necessitating rigorous auditing of both selection logic and justification language.
Statistical Analysis of Bias Types
The research employed a robust statistical framework to test for directional bias in recommendations. Convergence across four independent test families—binomial, chi-squared, Mann-Whitney U, and McNemar—demonstrates that the observed headline findings are not artifacts of any single test’s assumptions. Key comparisons included:
-
Flip Direction: The comparison between
Anti-Authority vs. Baseline
andContent-First vs. Baseline
was analyzed across various statistical metrics, yielding significant results (e.g., p < 0.001) for most tested conditions, indicating a strong pattern of bias manipulation based on instruction framing. -
Boost Direction: The analysis showed that while mild instructions can significantly cut the flip rate (e.g., by 12.9pp), they
barely move boost rate
(21.8% to 21.0%), suggesting that prompt-level debiasing is effective against swapped authority but not inflated authority signals, which are relevant toGEO-style citation-gaming.
Model Performance and Selection Agreement
When comparing model performance, the analysis of mean max h-index for picked papers revealed a hierarchy among the tested models (e.g., Mistral 48.7, Llama 3.1 45.5, Gemma 2 41.0). While establishing a strict population-level relationship between baseline preference and susceptibility is difficult with eight models, the extremes are aligned.
Furthermore, pairwise selection agreement analysis highlighted that the three frontier closed-weight models cluster tightly on selections:
-
gpt-5.4 claude-sonnet-4-6 agree on 63.0% of selections (the highest off-diagonal entry).
-
These three pairs are
consistent with a frontier post-training signature.
Drivers of Authority Bias and Justification Language
The gap between model capabilities is hypothesized to be driven by two primary factors: (i) the proportion of academic web text in pretraining corpora,
which dictates how strongly authority cues co-occur with quality judgments; and (ii) the post-training alignment regime
(e.g., RLHF), which modulates how heavily surface cues are weighted at decision time. This is further evidenced by the observed say-do gap,
where the 11.1pp difference suggests that instruction tuning allows models to readily say content-first while still selecting according to authority signals. In terms of justification patterns, models express authority bias through indirect language rather than naming specific signals:
-
The two most common categories are
Citations (citation count, highly cited)
andImpact (impactful, significant impact),
which together cover roughly two-thirds of mentions. -
Recency (14.3%) emerges as a third axis of non-content reasoning,
indicating that models frequently use publication date as a quality proxy.
Verbal vs. Behavioral Authority Mention Rates
A critical finding relates to the decoupling of verbal output from actual behavior. The frequency of authority mentions does not reliably predict susceptibility; for instance, claude-sonnet-4-6 talks about authority more than its behavior would suggest (17.5% mentions despite being the second most behaviorally resistant model).
This establishes that surface auditing alone is an unreliable proxy for behavioral bias even at the per-model level.
Improvements for AI systems
The findings presented in this paper are critically important because they empirically map the failure modes of current Large Language Models (LLMs)—specifically, their susceptibility to subtle authority cues, citation gaming, and metadata inflation. The key takeaway is that surface-level debiasing (e.g., be content-first
) is insufficient because models can say one thing while selecting based on another (the say-do gap
).
Given the high stakes of deploying these systems, I recommend implementing a multi-layered overhaul focusing on Alignment Refinement, Inference Architecture, and Evaluation Metrics.
The core problem is that current RLHF/Instruction Tuning often modulates surface generation tokens without correcting the underlying selection logits. We must modify the training objective to make selection accountable.
A. Authority-Aware Logit Modification (The Selection Guard
):
-
Improvement: Modify the Reward Model (RM) and fine-tuning objectives during RLHF. Instead of solely rewarding coherence or adherence to prompt instructions, introduce a penalty term that measures the divergence between the top K selected documents/citations and a truly balanced, content-driven selection set.
-
Mechanism: During training on citation-gaming datasets (like GEO-style), the RM must be trained not just on
Is this response good?
but also onDoes this response distribute its authority weight evenly across all credible sources, or does it disproportionately favor a single, highly cited source?
-
Goal: Force the model to learn that high citation count is a feature of the selection, not a determinant of quality. This addresses the finding that models can talk about authority but still select based on it.
B. Structured Citation Embedding (The De-Weighting
):
-
Improvement: Implement a structured embedding layer for citations that explicitly de-weights metadata signals during initial retrieval and selection phases, mimicking the
Content-First
instruction but at the architectural level. -
Mechanism: When generating a list of potential sources, the system should calculate an embedding score based on:
Score = alpha (Semantic Similarity) + beta (Novelty/Breadth) - gamma (Citation Count) eta
Where gamma is a dampening function (e.g., gamma=1 for the highest cited paper, 0.9 for the next, etc.) and eta > 1. This mathematically penalizes over-reliance on citation count or impact metrics during the initial candidate selection phase, making it harder to inflate authority via metadata.
These improvements modify how the model interacts with external knowledge bases at runtime, providing a robust defense layer.
A. Mandatory Source Triangulation:
-
Improvement: For any factual claim or significant conclusion, the system must be forced to retrieve and cite a minimum of three sources that support the claim, and these sources must come from different thematic areas or publication venues (e.g., one theoretical paper, one applied study, and one review article).
-
Function: If the model can only find supporting evidence heavily clustered around a single source (especially if that source is high-authority), the system must halt generation and output an uncertainty warning:
Warning: Supporting evidence is highly concentrated. Consider seeking corroboration from diverse fields.
B. Dual-Pass Verification (The Second Opinion
):
- Improvement: Implement a two-pass generation pipeline.
-
Pass 1 (Selection): The model selects the top N sources based purely on semantic relevance to the query, ignoring citation counts or prestige markers.
-
Pass 2 (Synthesis/Refinement): A secondary, smaller, and highly specialized LLM instance is prompted with the selected set of sources (not just one) and instructed:
Using ONLY these provided texts, synthesize an answer that explicitly attributes each claim to its source. Identify any contradictions or areas of disagreement.
- Benefit: This forces the model to demonstrate synthesis capability across multiple viewpoints rather than simply repeating the most authoritative-sounding paper.
Standard evaluation metrics are insufficient because they only measure if the answer is correct, not how it arrived there.
A. Authority Distribution Metric (ADM):
-
Improvement: Develop a quantitative metric to audit model outputs for authority bias. The ADM calculates the variance and entropy of the cited sources within an output's knowledge base footprint.
-
ADM = Entropy(Source Authority Scores) over sum Source Authority Scores
-
Goal: A high ADM (high entropy) indicates that the model has distributed its support across diverse and varied sources, suggesting robustness. A low ADM suggests over-reliance on a few high-authority sources, signaling potential bias or citation gaming susceptibility. This must be mandatory for all pre-deployment safety testing.
B. Adversarial Metadata Inflation
Benchmarking:
-
Improvement: Create a dedicated benchmark suite that systematically introduces fabricated authority signals (e.g., assigning fake
highly cited
tags or boosting the perceived recency of irrelevant papers) to test the model's resistance to these specific, non-semantic cues. -
Focus: Specifically test for the "Boosted pick > chance
and
Anti-Authority vs. Baseline (flip)" effects using controlled, simulated citation-gaming scenarios that are architecturally indistinguishable from real academic papers.
The resulting AI system will transition from being a sophisticated retriever to a demonstrably critical synthesis engine. It will not only answer the query but will also provide an audit trail of its own reasoning, quantifying the diversity and consensus among its supporting evidence.
The improved system can:
-
Resist Metadata Manipulation: It will ignore inflated signals (like citation count or prestige) unless those signals are consistently corroborated by multiple, semantically distinct sources.
-
Prove Synthesis Depth: It must synthesize answers from a minimum of three diverse viewpoints, explicitly flagging areas of disagreement rather than masking them under a single authoritative narrative.
-
Self-Audit Bias: It can provide an ADM score alongside its answer, alerting the user if the response is disproportionately reliant on a narrow cluster of high-authority sources, thereby enhancing user trust and mitigating legal/reputational risk associated with biased AI output.
Abstract
Large Language Models (LLMs) are increasingly used as conversational search engines for academic literature, yet whether they judge papers on content or on authority signals has not been tested causally. We investigate authority bias: systematic preference for papers based on author prestige, venue, and citations rather than content. Holding title and abstract constant, we vary authority metadata across three counterfactual conditions (original, flipped, boosted) over eight LLMs (five open-weight and three frontier closed-weight) in an in-context, single-turn, top-1 recommendation setting. Our experiments show that authority bias is substantial and directional, varies markedly across models, and is only partially addressable through prompt-level debiasing. We further document a say-do gap: debiasing instructions suppress authority mentions far faster than authority-driven flips, so surface auditing systematically underestimates behavioral bias.
Sources
- GPT-4 Technical Report
- Whose Name Comes Up? Auditing LLM-Based Scholar Recommendations
- GEO: Generative Engine Optimization
- Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias
- The Capacity for Moral Self-Correction in Large Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- A Survey on LLM-as-a-Judge
- Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection
- Mistral 7B
- In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Large Language Models as Recommender Systems: A Study of Popularity Bias
- Prestige over merit: An adapted audit of LLM bias in peer review
- Training language models to follow instructions with human feedback
- OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts
- C-SEO Bench: Does Conversational SEO Work?
- Qwen2.5 Technical Report
- Evaluating and Mitigating Discrimination in Language Model Decisions
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection