ReHoPER: Receding-Horizon Planning for Enhanced Reasoning
cs.CL, cs.AI, cs.LG
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/theSaeed/mobius
License: http://creativecommons.org/licenses/by/4.0/
The gist: We propose ReHoPER, an inference-only, zero-shot method that improves large language models' reasoning by generating and answering intermediate questions along multiple paths before the final answer.
Terminology
Abstract
We propose ReHoPER, an inference-only, zero-shot method that improves large language models' reasoning by generating and answering intermediate questions along multiple paths before the final answer. It iteratively plans a horizon of candidate intermediate questions, selects one to answer, and replans from the updated history. ReHoPER is task-agnostic, using the same generic instructions across datasets and models without labeled data or task-specific prompt design. Across multiple datasets, including iLLC, a new controlled benchmark for compositional reasoning, ReHoPER outperforms strong baselines, with the largest gains in the most compositional settings. Our implementation and the iLLC generator are publicly available to support future work.
Sources
- MoreHopQA: More Than Multi-hop Reasoning
- The Llama 3 Herd of Models
- Gemma 3 Technical Report
- Phi-4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen3 Technical Report
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering