Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

arXiv:2608.25660 · cs.CL, cs.AI · Submitted 2026-08-26 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-08-26

Updated: 2026-08-26

Comments: Accepted to EMNLP 2026 (Findings)

License: http://creativecommons.org/licenses/by/4.0/

The gist: Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas.

Terminology

Abstract

Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as "medium novel". To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent "medium novelty" bias.

Sources

Related papers