Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks
cs.AI
Submitted: 2026-08-30
Updated: 2026-08-30
Code: https://github.com/shitanshubhushan/Creativity-Evaluation-Framework
Terminology
Sources
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- AGC-Bench: Measuring Artificial General Creativity
- Divergent Creativity in Humans and Large Language Models
- MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
- CreativEval: Evaluating Creativity of LLM-Based Hardware Code Generation
- MMTEB: Massive Multilingual Text Embedding Benchmark
- CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity
- AIDE: AI-Driven Exploration in the Space of Code
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations
- EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
- Kosmos: An AI Scientist for Autonomous Discovery
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
- OpenAI GPT-5 System Card
- GLM-5: from Vibe Coding to Agentic Engineering
- Qwen3 Technical Report
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection