AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/AutoDataBench/AutoDataBench
Terminology
Sources
- Program Synthesis with Large Language Models
- FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
- LoRA: Low-Rank Adaptation of Large Language Models
- DCA-Bench: A Benchmark for Dataset Curation Agents
- Qwen2.5-Coder Technical Report
- Can Generalist Agents Automate Data Curation?
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Propagating Knowledge Updates to LMs Through Distillation
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Representation Learning with Contrastive Predictive Coding
- MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection
- FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- Qwen2 Technical Report
- Qwen3 Technical Report
- SWE-smith: Scaling Data for Software Engineering Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering