Think Thrice Before Reranking: Multi-perspective Evidence and Reasoning Integration for Text Reranking
cs.IR, cs.CL
Submitted: 2026-09-17
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reasoning-based reranking with Large Language Models (LLMs) has shown promising improvements in text ranking.
Terminology
Abstract
Reasoning-based reranking with Large Language Models (LLMs) has shown promising improvements in text ranking. However, current methods predominantly rely on a single reasoning trajectory, resulting in rankings that are susceptible to reasoning errors and inherently constrained in modeling the multifaceted signals underlying document relevance. To resolve this dilemma, we propose MERIT-Rank(Multi-perspective Evidence and Reasoning Integration for Text Reranking), a framework that models complementary reasoning trajectories to improve reranking robustness. MERIT-Rank formulates a Multi-Trajectory Reasoning Space (MTRS) that evaluates query-document relevance from multiple perspectives and introduces a joint reranker that consolidates these reasoning paths into a unified ranking decision. We further develop Progressive Rank Policy Optimization (PRPO), a progressive training framework that stabilizes reasoning trajectories while continually improving ranking quality through staged optimization objectives. Experiments on both reasoning-intensive and traditional retrieval benchmarks show that MERIT-Rank consistently achieves superior performance over competitive baselines. The 4B model notably outperforms most 7B and even 32B rerankers on BRIGHT.
Sources
- Overview of the TREC 2020 deep learning track
- Overview of the TREC 2019 deep learning track
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Holistic Evaluation of Language Models
- ReasonRank: Empowering Passage Ranking with Strong Reasoning Ability
- GPT-4 Technical Report
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- DemoRank: Selecting Effective Demonstrations for Large Language Models in Ranking Task
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents
- Rank1: Test-Time Compute for Reranking in Information Retrieval
- Rank-K: Test-Time Reasoning for Listwise Reranking
- Zero-Shot Listwise Document Reranking with a Large Language Model
- JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
- Multi-Stage Document Ranking with BERT
- The Expando-Mono-Duo Design Pattern for Text Ranking with Pretrained Sequence-to-Sequence Models
- RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze!
- Improving Passage Retrieval with Zero-Shot Question Generation
- InsertRank: LLMs can reason over BM25 scores to Improve Listwise Reranking
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG