Can Agents Design Libraries for Agents?
cs.AI, cs.CL, cs.SE
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/SprocketLab/librarydesignbench
Terminology
Sources
- Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
- Code for Machines, Not Just Humans: Quantifying AI-Friendliness with Code Health Metrics
- Large Language Models as Tool Makers
- When LLMs Meet API Documentation: Can Retrieval Augmentation Aid Code Generation Just as It Helps Developers?
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- GLM-5: from Vibe Coding to Agentic Engineering
- LILO: Learning Interpretable Libraries by Compressing and Documenting Code
- Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
- ADK Arena: Evaluating Agent Development Kits via LLM-as-a-Developer
- On Mitigating Code LLM Hallucinations with API Documentation
- Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents
- Kimi K3: Open Frontier Intelligence
- Refactoring Codebases through Library Design
- DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation
- ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
- Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
- A Survey on Large Language Model Impact on Software Evolvability and Maintainability: the Good, the Bad, the Ugly, and the Remedy
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection