IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR
cs.CL, cs.AI
Submitted: 2026-01-23
Updated: 2026-03-06
Comments: 24 Pages, v2, Abstract Modified
DOI: 10.18653/v1/2026.findings-acl.1256
License: http://creativecommons.org/licenses/by/4.0/
The gist: Peer review relies on substantive, evidence-based questions, yet current LLMs generate surface-level queries that perform worse than human reviewer questions in expert evaluation.
Terminology
Abstract
Peer review relies on substantive, evidence-based questions, yet current LLMs generate surface-level queries that perform worse than human reviewer questions in expert evaluation. To address this gap, we curate a high-quality dataset of reviewer questions from OpenReview and conduct a human preference study where expert annotators evaluate question-paper pairs across three dimensions: effort, evidence, and grounding. From these annotations, we train IntelliReward, a reward model built from a frozen autoregressive LLM with trainable multi-head transformers. Validated against expert judgments, IntelliReward predicts reviewer-question quality better than API-based SFT baselines and provides scalable evaluation. We apply Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) with IntelliReward to train IntelliAsk, a question-generation model aligned with human standards of effortful, evidence-based critique. Human evaluations show IntelliAsk generates more grounded, substantive and effortful questions than strong baselines and reduces reliance on first-page content. We also find improvements on reasoning and writing benchmarks, suggesting reviewer-question quality correlates with broader capabilities. Compared to Qwen3-32B, IntelliAsk improves MuSR (68.3 vs 64.7 Acc) and WritingBench (8.31 vs 8.07). We release our code, filtered review dataset, expert annotations, IntelliAsk and IntelliReward to support automatic evaluation of grounding, effort, and evidence in LLM-generated review questions.
Sources
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- Graph-Guided Passage Retrieval for Author-Centric Structured Feedback
- WebGPT: Browser-assisted question-answering with human feedback
- MARG: Multi-Agent Review Generation for Scientific Papers
- A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
- olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- WritingBench: A Comprehensive Benchmark for Generative Writing
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering