Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data
cs.SE, cs.AI
Submitted: 2026-07-20
Updated: 2026-08-27
Code: https://github.com/karpathy/autoresearch
Terminology
Sources
- Smartajweed Automatic Recognition of Arabic Quranic Recitation Rules
- Concrete Problems in AI Safety
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Categorizing Variants of Goodhart's Law
- MLGym: A New Framework and Benchmark for Advancing AI Research Agents
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties