PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 30 pages, 5 figures, 16 tables. Code and data: https://github.com/Henri-XYu02/PrimeScientist
Code: https://github.com/Henri-XYu02/PrimeScientist
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- MARS: Modular Agent with Reflective Search for Automated AI Research
- Reasoning Models Don't Always Say What They Think
- SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning
- Rational Metareasoning for Large Language Models
- Token-Budget-Aware LLM Reasoning
- AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs
- FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
- Pancake: Hierarchical Memory System for Multi-Agent LLM Serving
- TritonDFT: Automating DFT with a Multi-Agent Framework
- GEAR: Genetic AutoResearch for Agentic Code Evolution
- AIDE: AI-Driven Exploration in the Space of Code
- GPT is becoming a Turing machine: Here are some ways to program it
- Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Bilevel Autoresearch: Meta-Autoresearching Itself
- Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings
- Effort Allocation for Deadline-Aware Task and Motion Planning: A Metareasoning Approach
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
- DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering