TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
cs.LG, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-27
Code: https://github.com/openai/codex
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over
Terminology
Abstract
Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competitions still finishes below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4,465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect. Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold instead collapses into a narrow loop: Codex spends its steps re-weighting ensembles and tuning submissions, MLEvolve mutates its model in place, and neither pivots at the human rate nor reopens abandoned work. A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions. We release the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.
Sources
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery
- Kolb-Based Experiential Learning for Generalist Agents with Human-Level Kaggle Data Science Performance
- AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
- HCAST: Human-Calibrated Autonomy Software Tasks
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
- The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks