LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation
Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li, Yujie Zhang, Zichen Zhu, Zhiying Zou, Kai Yu, Lu Chen
Shanghai Jiao Tong University · Shanghai Innovation Institution
cs.CL, cs.AI, cs.DB, cs.MA
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 17 pages
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation Abstract With the rapid advancement of large language models (LLMs), research idea generation has attracted
Terminology
Summary
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation
Abstract
With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability to provide unified and reliable assessments across a coherent distribution of generated ideas. To address this challenge, we propose LigBench, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions. In addition, we introduce PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation. Extensive experiments demonstrate that LigBench achieves stable and interpretable evaluations, significantly improving alignment with expert judgments. Furthermore, models trained on PAIR-IQ exhibit enhanced ranking accuracy and robustness, establishing a principled standard for scalable and objective research idea assessment.
Introduction
The rapid evolution of LLMs has ushered in a new era of automated research idea generation. Building on the strong reasoning and knowledge synthesis capabilities of these models, a growing body of recent work has proposed LLM-based frameworks for scientific ideation, enabling the generation of novel hypotheses, methodological variations, and promising research directions grounded in large-scale scientific literature. Such capabilities hold great promise for accelerating discovery across disciplines, guiding researchers toward unexplored directions, and reducing the time and cost associated with manual ideation. However, as the utility of these systems continues to grow, the lack of rigorous, objective, and reproducible evaluation methodologies remains a critical bottleneck for the systematic assessment of generated ideas.
Current evaluation strategies for idea generation are often fragmented. Some studies rely on direct scoring by LLMs, where models assign quality scores to their own or peer-generated ideas, while others depend on human experts to perform pairwise or ranking-based evaluations. Automated metrics such as novelty, diversity, or semantic similarity offer partial insights but fail to capture the multi-dimensional nature of research idea quality. Overall, these approaches lack consistency: they may be sensitive to prompt phrasing, suffer from inter-annotator variability, or provide only relative assessments. Consequently, there is no unified framework to benchmark idea generation systems comprehensively.
Pairwise comparison has emerged as a promising approach to sidestep some subjective biases inherent in absolute scoring. In this paradigm, two candidate ideas are presented side by side, and an evaluator (whether human or automated) selects the one deemed superior along specified criteria. Aggregating multiple pairwise comparisons can yield a ranking of ideas with higher consistency than solitary ratings. Nevertheless, pairwise evaluation alone does not constitute a full benchmarking solution: it requires a sufficiently large and diverse pool of test pairs, clear annotation protocols, and scalable tooling to collect judgments at scale.
To address these challenges, we introduce LigBench, an automated evaluation framework tailored for research idea assessment. LigBench operationalizes multi-dimensional quality criteria by orchestrating large-scale pairwise comparisons, metadata tracking, statistical aggregation, and automatic score updates into a cohesive pipeline. We drew on the thinking process experts use when evaluating an idea. Experts typically begin by searching for existing results related to the idea for reference or comparison. In addition, pairwise comparison alone is not sufficient to objectively and convincingly reflect the quality of an idea. Therefore, we have designed and provided a complete process based on pairwise comparison combined with the Elo mechanism to continuously update scores. We also present the PAIR-IQ dataset, a curated dataset of research papers' ideas containing structured idea representations. PAIR-IQ serves both as a gold-standard reference for comparative evaluation and as a training resource for models designed to predict pairwise judgments.
Main Contributions
-
Unified Evaluation Framework: We introduce LigBench as a unified and extensible evaluation framework that systematically integrates data curation, idea retrieval, pairwise comparison, score propagation, and result aggregation within a single automated pipeline.
-
PAIR-IQ Dataset: We release PAIR-IQ dataset, comprising over 11000 conference papers with debiased evaluation scores and formalized representations to support objective assessment of research ideas.
-
Open-Source Resources: We will release the PAIR-IQ dataset together with the complete LigBench evaluation pipeline to facilitate reproducibility and enable community-driven extensions. The PAIR-IQ dataset is publicly available at https://huggingface.co/datasets/USER3IjEBHj9/PAIR-IQ/tree/main. The full LigBench evaluation pipeline will be released in the near future.
-
Benchmarking and Analysis: We conduct extensive experiments with state-of-the-art LLMs, demonstrating that LigBench can discern subtle differences in idea quality and offering insights into model strengths and weaknesses.
Method
To provide a unified, objective, and human-aligned evaluation of research ideas, we propose LigBench, an automated benchmark designed for systematic idea assessment. The evaluation of each idea is decomposed into four aspects: rating, contribution, soundness, and novelty, with each aspect scored on a 0-5 scale. Here, rating represents the overall quality as perceived by reviewers, contribution reflects the significance of the idea within its research domain, soundness captures methodological rigor and validity, and novelty measures the originality of the proposed idea.
Following the pipeline, each idea is first transformed into a formalized representation and then iteratively compared against similar papers retrieved from the database using a large language model. In each comparison, the model evaluates the relative quality between idea pairs, and the resulting judgments are aggregated with the objective scores from the PAIR-IQ dataset to update the score of the target idea. This process is repeated until convergence, producing a final score that is consistent across different sources and aligned with human preferences.
Idea Formalization
A fundamental challenge in building a unified idea evaluation framework is the heterogeneity in the formats of the ideas under evaluation. Due to inherent differences in the distributions of ideas generated by different models and frameworks, directly comparing these ideas is prone to bias. For instance, when using large language models as evaluators, they often tend to favor ideas that are described in greater detail, potentially overlooking those that are more concise yet fundamentally more innovative or valuable. Moreover, different studies are conducted under diverse frameworks, leading to substantial variations in how final ideas are presented, including their structure and formatting, which further complicates fair and consistent comparison across models and methods.
To mitigate these issues, it is necessary to standardize ideas across different distributions. The framework focuses on the core aspects and boundaries of each idea and decouples it into four distinct components:
-
Main Target: a concise, one-sentence summary of the task and objective that the idea aims to address.
-
Core Breakthrough: the key advancement or novel contribution introduced by the idea compared to prior work.
-
Innovative Methods: the concrete methodological approaches or techniques proposed to achieve the target.
-
Experimental Design: the experimental setup and evaluation plan designed to validate the effectiveness of the proposed methods.
Formally, let I ∈ I denote an input idea from the idea space. We define an LLM-based decomposition operator DLLM: I → T × B × M × E such that (T, B, M, E) = DLLM (I), where T ∈ T represents the main target, B ∈ B denotes the core breakthrough, M ∈ M corresponds to the innovative methods, and E ∈ E specifies the experimental design. This formulation casts idea formalization as a structured, LLM-driven mapping into a unified representation space, facilitating systematic comparison and subsequent quantitative evaluation.
By explicitly disentangling complementary semantic facets of an idea into structured components, this LLM-driven decomposition yields a unified yet fine-grained representation that captures the substantive content while abstracting away superficial variations in expression and format. Consequently, it mitigates evaluation bias arising from heterogeneous idea sources and presentation styles, enabling more objective and consistent comparison across models and frameworks. Furthermore, this structured representation establishes a principled foundation for subsequent quantitative analyses, including assessments of soundness, novelty, contribution, and other relevant criteria.
PAIR-IQ Construct
To better evaluate research ideas according to human preferences, we collected and constructed the PAIR-IQ dataset, comprising over 11,000 papers from ICLR 2024, ICLR 2025, and NeurIPS 2024, including oral, spotlight, poster, and rejected papers. For each paper, we retrieved ratings, contribution scores, and soundness scores from OpenReview, and projected each score to a standardized 0-5 range to ensure comparability across different venues and scoring scales.
The dataset draws from three major machine learning venues, with ICLR 2025 contributing the largest proportion (36.7%), followed by ICLR 2024 (32.4%) and NeurIPS 2024 (30.9%). The distribution across acceptance categories shows that poster papers constitute the majority (57.3%), followed by rejected submissions (36.6%), spotlight papers (4.3%), and oral presentations (1.8%). This distribution ensures that the dataset captures a wide spectrum of paper quality levels, from top-tier accepted works to borderline and rejected submissions. The dataset spans 12 research areas, with large language models (LLM) representing the most prevalent topic (19.3%), followed by training methodologies (15.4%) and reinforcement learning (13.4%), reflecting the current research trends in the machine learning community.
For each paper, we parsed all textual content and employed an LLM to extract structured components, including Main Target, Core Breakthrough, Innovative Methods, and Experimental Design. These components provide a standardized representation of each research idea, enabling consistent downstream evaluation and facilitating subsequent scoring procedures.
Next, we address potential biases in the collected ratings. Since different venues and years may exhibit systematic differences in reviewer behavior, the resulting score distributions can vary significantly across the dataset. To ensure that the scores of all papers are comparable and the overall distribution is consistent, we perform score debiasing using a mean-shifting procedure. Specifically, let s i(v) denote the original score (rating, contribution, or soundness) of paper i from venue v, with venue mean μ v and overall mean μ all. The debiased score is computed as ŝ i(v) = s i(v) − μ v + μ all. This procedure aligns the score distributions across different venues and years, ensuring that subsequent evaluations reflect the intrinsic quality of each idea rather than systematic biases in reviewer ratings. Additionally, the dataset was further processed through embedding computation to facilitate downstream evaluation tasks.
Pairwise Comparison
When a new research idea is submitted for evaluation, we first transform it into a formalized representation using the same procedure described previously. This ensures that the target idea and reference papers are represented in a unified format.
Based on this representation, we retrieve semantically related papers from the database using a parallel retrieval strategy that combines keyword matching and embedding-based similarity. The target idea is then compared pairwise with each retrieved paper using an LLM, which assesses relative quality along three dimensions: rating, contribution, and soundness.
We update the target idea's scores using an adapted Elo rating algorithm. In standard Elo, the expected probability that idea A wins over idea B is: E A = 1 / (1 + 10(s B − s A)/d), where s A and s B denote the current scores of ideas A and B, and d is a scaling parameter. Given the observed outcome S A ∈ 0, 0.5, 1, the score is updated as: s new A = C(s A + K · (S A − E A)), where K is an adaptive update factor and C(·) is a soft clamping function that constrains scores to [0, 5].
Through multiple rounds of pairwise comparisons, the target idea's scores gradually converge. The evaluation terminates when score changes fall below a predefined threshold, yielding the final assessment.
Novelty Assessment
Through the above procedures, we obtain objective evaluations of research ideas in terms of rating, contribution, and soundness. However, these dimensions do not explicitly capture the novelty of an idea. Unlike other evaluation dimensions that can be assessed through comparative analysis, novelty requires determining whether an idea introduces genuinely new insights compared to existing literature.
To address this challenge, we design a novelty assessment module that combines LLM-based initial estimation with similarity-based quantitative analysis. The evaluation proceeds in three stages: (1) obtaining an initial novelty estimate s novelty(0) from the LLM evaluator; (2) retrieving semantically related papers via the Semantic Scholar API and computing a weighted similarity metric s weighted, which is then mapped to a novelty score s sim through an inverse sigmoid function; and (3) fusing both assessments into the final score: s final novelty = β · s sim + (1 − β) · s novelty(0), where β = 0.7 assigns higher weight to the similarity-based calculation, reflecting its objectivity, while retaining the LLM's judgment to capture aspects not fully represented by semantic similarity alone.
Experiments
Pairwise Evaluation
The accuracy of pairwise judgments made by the LLM serves as a key guarantee for the reliability of the entire evaluation framework. To assess this reliability, we conduct a dedicated pairwise evaluation on the three dimensions: rating, contribution, and soundness. For this purpose, we randomly sample a held-out test set consisting of 269 pairs of thematically similar papers from the database, ensuring that none of the selected papers appear in the training data.
The ground-truth labels for pairwise comparison are derived from the debiased OpenReview scores. The accuracy of the LLMs pairwise judgments is then measured by comparing its predictions against these ground-truth relationships. Overall, stronger models consistently achieve higher accuracy across all three dimensions, indicating that pairwise judgment is a non-trivial capability that benefits from increased model capacity and reasoning ability. Among the off-the-shelf models, GPT-5 substantially outperforms other baselines, achieving accuracies above 0.80 on both rating and soundness, which suggests a high level of agreement with debiased human judgments. In contrast, smaller open-source models such as Qwen2.5-7B and Qwen2.5-14B exhibit limited performance, with accuracies close to random guessing in some dimensions, particularly on rating and soundness.
Notably, models trained on the PAIR-IQ dataset show substantial improvements across all dimensions. Both trained variants achieve consistently higher accuracies than their untrained counterparts, demonstrating that targeted pairwise supervision effectively enhances the models ability to align with human evaluation criteria. These results validate the design of the PAIR-IQ dataset and further support the use of pairwise judgment as a reliable component within the LigBench evaluation framework.
Although the pairwise judgment accuracy does not reach near-perfect levels, this outcome also reflects the intrinsic difficulty of the task, as distinguishing fine-grained differences between closely related research ideas is inherently challenging. Nevertheless, when applied within the LigBench framework, even imperfect pairwise judgments can still lead to reliable final evaluations. Starting from a strong base model, repeated score updates over a large number of pairwise comparisons allow the target scores to gradually converge toward reasonable values. Further analysis of the error cases reveals that most incorrect judgments occur when the debiased scores of the two compared papers are very close, meaning that even when the LLM makes an incorrect decision, the resulting update magnitude is limited.
Human Alignment
Although the pairwise judgment capability of the large language model has been validated using debiased OpenReview scores, it remains essential to perform human verification. To further verify human alignment, we recruited several PhD-level AI researchers to conduct pairwise comparisons on 100 randomly sampled idea pairs. For each pair, experts selected the superior idea along the three dimensions: rating, contribution, and soundness. The resulting human judgments were then compared with the LLMs pairwise predictions. The LLM exhibits substantial agreement with human experts across all three dimensions. In particular, the highest alignment is observed for contribution, indicating that the model is effective at assessing the relative significance of research ideas within a domain. The consistently strong agreement on rating and soundness further suggests that the LLMs pairwise judgments capture key aspects of overall quality and methodological rigor.
Paper Evaluation
We conducted an evaluation on papers from NeurIPS 2025 to demonstrate the effectiveness of our framework. The evaluation applied the full LigBench procedure. Papers that were accepted at NeurIPS 2025 consistently receive higher scores across all four evaluation dimensions compared to rejected papers. The largest differences are observed in novelty and rating, suggesting that both originality and overall perceived quality are strong determinants of acceptance. Differences in contribution and soundness are also notable, indicating that methodological rigor and domain impact are effectively captured by LigBench. These results demonstrate that LigBench can meaningfully distinguish higher-quality research ideas from lower-quality ones, even when applied to a relatively small test set of 50 papers.
Impact of Data Source and Debiasing
To examine the consistency of our evaluation framework across different data sources, we analyze the impact of the three major sources in the PAIR-IQ dataset: ICLR 2024, ICLR 2025, and NeurIPS 2024. We conduct an ablation study where the score updating process is performed using only a single source at a time. The differences between the scores obtained using a single source and those obtained using all sources together are minor, with all metrics showing deviations within a few percentage points. This indicates that our evaluation results are largely insensitive to the choice of data source. Consequently, the debiasing strategy effectively mitigates distributional differences across venues, enabling stable and comparable evaluation regardless of whether a single source or all sources are used.
Benchmark and Model Evaluation
To comprehensively assess the quality of LLM-generated research ideas, we evaluate both standalone LLMs and idea-generation frameworks using LigBench. Each system is prompted to freely generate 10 ideas across 11 research topics, including Reinforcement Learning, Image Generation, AI Safety, Agent, Computer Vision, Audio and Speech, Embodied Intelligence, Multimodal, AI for Science, Large Language Models, and Training. Final scores are computed as the average across all generated ideas.
We compare two representative idea-generation frameworks, CoI and SciPIP, against direct LLM baselines. Both employ GPT-5 as their backbone, enabling a controlled comparison that isolates the effect of framework design. Among standalone LLMs, GPT-5.2 and GPT-5 clearly outperform GPT-4.1 and GPT-4o across all dimensions, with GPT-4o scoring below 1.0 on rating and contribution, indicating that earlier models struggle to meet basic quality thresholds. Idea-generation frameworks do not consistently outperform direct prompting on strong backbone models. Despite using GPT-5, CoI and SciPIP often score lower than standalone GPT-5, suggesting that these frameworks may optimize weaker models but introduce constraints for capable ones. Notably, SciPIP achieves the highest Soundness score among all evaluated systems, slightly exceeding its underlying model GPT-5. Finally, while novelty scores broadly follow expectations, absolute differences are small, highlighting that generating truly novel ideas remains challenging for current LLMs.
Conclusion
We introduced LigBench, a unified and human-aligned benchmark for evaluating research ideas generated by large language models and agentic frameworks. By decomposing idea quality into four dimensions—rating, contribution, soundness, and novelty—LigBench enables consistent and fine-grained assessment across diverse sources. To support reliable evaluation, we constructed the PAIR-IQ dataset with standardized, debiased scores, enabling effective pairwise judgment and robust score updating. Results show that LigBench produces stable evaluations aligned with expert preferences, even when individual pairwise judgments are imperfect. LigBench provides a practical foundation for benchmarking future idea-generation systems and advancing automated evaluation of scientific creativity. It can also serve as a reliable reward signal for training idea-generation models via reinforcement learning, enabling more scalable and human-aligned AI-driven idea generation.
Improvements for AI systems
Improvements to AI Systems Based on LigBench:
-
Human-Aligned Idea Evaluation Module: Integrate LigBench's four-dimensional scoring (rating, contribution, soundness, novelty) into AI systems as a built-in reward function. This enables AI to self-assess generated research ideas against human expert standards, replacing arbitrary self-scoring with calibrated, debiased metrics.
-
Structured Idea Formalization Capability: Implement the LLM-driven decomposition operator (D LLM) that breaks any research idea into Main Target, Core Breakthrough, Innovative Methods, and Experimental Design. This allows AI systems to standardize heterogeneous idea formats, reducing evaluation bias from verbosity or presentation style, and enabling fair cross-framework comparisons.
-
Pairwise Judgment Training with PAIR-IQ: Fine-tune AI models on the PAIR-IQ dataset (11,000+ papers with debiased scores) to improve pairwise comparison accuracy. The improved AI can reliably judge which of two research ideas is superior across rating, contribution, and soundness, achieving up to 80%+ agreement with human experts—critical for ranking and selection tasks.
-
Elo-Based Iterative Score Refinement: Equip AI systems with the adapted Elo rating algorithm (with soft clamping to [0,5]) to continuously update idea scores through repeated pairwise comparisons. This enables AI to converge on stable, objective quality scores even when individual judgments are imperfect, mimicking expert review processes.
-
Hybrid Novelty Assessment: Incorporate the novelty module that fuses LLM-based initial estimation (30% weight) with Semantic Scholar similarity-based quantitative analysis (70% weight via inverse sigmoid mapping). This allows AI to objectively measure originality against existing literature, reducing hallucinated novelty claims and grounding assessments in real-world prior work.
-
Debiased Evaluation Across Venues: Apply the mean-shifting debiasing procedure to normalize scores across different conferences, years, and reviewer behaviors. AI systems can use this to evaluate ideas from diverse sources without systematic bias, ensuring fair comparisons regardless of origin.
-
Retrieval-Augmented Comparison Pipeline: Implement the parallel retrieval strategy (keyword + embedding similarity) to fetch relevant papers from a curated database. The improved AI can ground its evaluations in real, comparable prior work, enhancing contextual understanding and reducing reliance on internal knowledge alone.
What the Improved AI System Can Do:
-
Generate and Self-Evaluate Research Ideas: Propose novel hypotheses and simultaneously score them on four expert-aligned dimensions, providing immediate, objective feedback without human intervention.
-
Rank and Select Best Ideas: Automatically compare thousands of candidate ideas (e.g., from different models or frameworks) using pairwise judgments and Elo scoring, identifying top-tier contributions with human-level accuracy.
-
Provide Explainable Quality Assessments: Output structured scores (e.g.,
rating: 3.8, contribution: 4.2, soundness: 3.5, novelty: 4.0
) with justification grounded in comparisons to specific retrieved papers, making AI reasoning transparent. -
Train via Reinforcement Learning: Use LigBench's scores as a reliable reward signal for RL fine-tuning, enabling AI to iteratively improve its idea generation toward human-preferred outcomes at scale.
-
Benchmark Competing AI Systems: Serve as a standardized evaluator to compare different LLMs or idea-generation frameworks (e.g., CoI vs. SciPIP) on equal footing, revealing which systems genuinely advance scientific creativity.
-
Filter Low-Quality Submissions: Automatically reject or flag weak research proposals early (e.g., scoring below 1.0 on rating), saving human reviewers time and ensuring only high-potential ideas proceed.
Abstract
With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability to provide unified and reliable assessments across a coherent distribution of generated ideas. To address this challenge, we propose LigBench, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions. In addition, we introduce PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation. Extensive experiments demonstrate that LigBench achieves stable and interpretable evaluations, significantly improving alignment with expert judgments. Furthermore, models trained on PAIR-IQ exhibit enhanced ranking accuracy and robustness, establishing a principled standard for scalable and objective research idea assessment.
Sources
- GPT-4 Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Distilling the Knowledge in a Neural Network
- Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas
- The Semantic Scholar Open Data Platform
- Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- AI Idea Bench 2025: AI Research Idea Generation Benchmark
- Qwen2.5 Technical Report
- Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- SciPIP: An LLM-based Scientific Paper Idea Proposer
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Qwen3 Technical Report
- Hypothesis Generation with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering