Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research
cs.SE, cs.AI
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/Flavorfish/AutoRepro
Terminology
Sources
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
- From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery
- Claw AI Lab: An Autonomous Multi-Agent Research Team
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties