Diagnosing LLM Reranker Behavior Under Fixed Evidence Pools
cs.LG, cs.CL, cs.IR
Submitted: 2026-02-20
Updated: 2026-09-08
Comments: Updated to the version published at SIGIR 2026
Journal ref: Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), pp. 3587-3592, 2026
Code: https://github.com/barisarat/llm_reranker_multinews
License: http://creativecommons.org/licenses/by/4.0/
The gist: Standard reranking evaluations study how a reranker orders candidates returned by an upstream retriever.
Terminology
Abstract
Standard reranking evaluations study how a reranker orders candidates returned by an upstream retriever. This setup couples ranking behavior with retrieval quality, so differences in output cannot be attributed to the ranking policy alone. We introduce a controlled diagnostic for reranking that uses Multi-News clusters as fixed evidence pools. We limit each pool to eight documents and pass identical inputs to all rankers. Within this setup, BM25 and MMR serve as interpretable reference points for lexical matching and diversity optimization. Across 345 clusters, we find that redundancy patterns vary by model: one LLM implicitly diversifies at larger selection budgets, while another increases redundancy. In contrast, LLMs underperform on lexical coverage at small selection budgets. As a result, LLM rankings diverge substantially from both baselines rather than consistently approximating either strategy. By reducing retrieval variance through fixed pools, we interpret these differences more directly as differences in ranking policy. This diagnostic is model agnostic and can be applied to any ranker, including open source systems and proprietary APIs. Our code and processed data for both the Multi-News diagnostic and the complementary TREC-DL evaluation are publicly available at https://github.com/barisarat/llm reranker multinews.git.
Sources
- The Llama 3 Herd of Models
- Language Models (Mostly) Know What They Know
- Holistic Evaluation of Language Models
- Zero-Shot Listwise Document Reranking with a Large Language Model
- Passage Re-ranking with BERT
- Document Ranking with a Pretrained Sequence-to-Sequence Model
- RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models
- Qwen2.5 Technical Report
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- RankLLM: A Python Package for Reranking with LLMs
- OpenAI GPT-5 System Card
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks