Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers
cs.CL, cs.IR, cs.LG
Submitted: 2026-08-27
Updated: 2026-09-26
Comments: 9 pages main text, 45 pages total
Code: https://github.com/thomsonreuters/presentation-dependence
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- One Pass, Any Order: Position-Invariant Listwise Reranking for LLM-Based Recommendation
- Order Independence With Finetuning
- TourRank: Utilizing Large Language Models for Documents Ranking with a Tournament-Inspired Strategy
- Overview of the TREC 2019 deep learning track
- From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective
- How to Evaluate Reward Models for RLHF
- Gemma 4 Technical Report
- Large Language Models are Zero-Shot Rankers for Recommender Systems
- LoRA: Low-Rank Adaptation of Large Language Models
- Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking
- Don't "Overthink" Passage Reranking: Is Reasoning Truly Necessary?
- ProRank: Prompt Warmup via Reinforcement Learning for Small Language Models Reranking
- R-Drop: Regularized Dropout for Neural Networks
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
- RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
- Decoupled Weight Decay Regularization
- Rethinking Reasoning in Document Ranking: Why Chain-of-Thought Falls Short
- Learning from Emptiness: De-biasing Listwise Rerankers with Content-Agnostic Probability Calibration
- Fine-Tuning LLaMA for Multi-Stage Text Retrieval
- RewardBench 2: Advancing Reward Model Evaluation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering