Measuring the Creativity of Frontier LLMs in Automated Research
cs.CL
Submitted: 2026-09-12
Updated: 2026-10-07
License: http://creativecommons.org/licenses/by/4.0/
The gist: Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated.
Terminology
Abstract
Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. In this paper, we propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty. Valueness assesses whether each proposed idea is useful, while novelty is evaluated from three perspectives: whether the same idea has appeared before (Exact-Match P-Novelty), whether a previously unexplored variable or variable combination is explored (Variable-level P-Novelty), and whether the idea directly follows retrieved external knowledge or departs from it (H-Novelty). Our evaluation shows that the models achieve relatively similar scores on most creativity metrics, but differ substantially in Variable-level P-Novelty, which reflects the breadth of research-space exploration. Further correlation and idea-level performance analyses show that Variable-level P-Novelty is the creativity dimension most strongly associated with research performance.
Sources
- Accelerating scientific discovery with Co-Scientist
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Unlocking LLM Creativity in Science through Analogical Reasoning
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering