Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation
cs.CL
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/allenai/asta-paper-finder
Terminology
Sources
- ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
- Measuring the Gap Between Human and LLM Research Ideas
- Accelerating scientific discovery with Co-Scientist
- All That Glitters is Not Novel: Plagiarism in AI Generated Research
- Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas
- Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents
- Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications
- An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics
- ScholarEval: Research Idea Evaluation Grounded in Literature
- InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
- Human-LLM Compound System for Scientific Ideation through Facet Recombination and Novelty Evaluation
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
- Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas
- Literature-Grounded Novelty Assessment of Scientific Ideas
- Unlocking LLM Creativity in Science through Analogical Reasoning
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
- The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
- Large Language Models are not Fair Evaluators
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering