AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
Jia Liu, Veena Krishnaraj, Kateryna Vovk, Kosuke Aizawa, Adrian E. Bayer, Linda Blot, Jessica Cowell, Suyog Garg, Jonathan Grée, Anamaria Hell, Ben Horowitz, Masaya Ichikawa, Kanyuni Iemoto, Keigo Kondo, Zacharie Lorsin, Kevin McCarthy, Jamie Robinson, Miguel Ruiz-Granda, Leander Thiele, Ievgen Vovk, Mingshen Zhou
cs.CL, astro-ph.CO, astro-ph.IM, cs.HC, gr-qc
Submitted: 2026-07-28
Comments: 16 pages, 4 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- AstroLLaMA: Towards Specialized Foundation Models in Astronomy
- Mitigating the Bias of Large Language Model Evaluation
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- LLM Evaluators Recognize and Favor Their Own Generations
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- Self-Preference Bias in LLM-as-a-Judge
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- Can AI-Generated Text be Reliably Detected?
- Assisting Research Proposal Writing with Large Language Models: Evaluation and Refinement
- Can Large Language Models Evaluate Grant Proposal Quality? Revisiting the Wenner{\aa}s and Wold Peer Review Data
- Evaluating LLM-Based Grant Proposal Review via Structured Perturbations
- The Denario project: Deep knowledge AI agents for scientific discovery
- LLMs with in-context learning for Algorithmic Theoretical Physics
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering