Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty
cs.CL, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: Accepted to EMNLP 2026 (Findings)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas.
Terminology
Abstract
Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as "medium novel". To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent "medium novelty" bias.
Sources
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- Artificial Intelligence and Natural Language Processing and Understanding in Space: A Methodological Framework and Four ESA Case Studies
- MemoNoveltyAgent: A Historical Research Memory-Aware Agent Workflow for Paper Novelty Assessment
- Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents
- Harnessing Large Language Models for Scientific Novelty Detection
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Are We Truly Innovating? A Qualitative and Quantitative Study of Originality in AI Research Papers
- gpt-oss-120b & gpt-oss-20b Model Card
- OpenAI GPT-5 System Card
- SciPIP: An LLM-based Scientific Paper Idea Proposer
- NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment
- Qwen3 Technical Report
- Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering