Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
cs.SE, cs.AI, cs.CL
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/haw-ai-i/roundtrip
Terminology
Sources
- Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
- Program Synthesis with Large Language Models
- SWE-Bench 5G: Benchmarking AI Coding Agents on Telecom Network Engineering Tasks
- Evaluating Large Language Models Trained on Code
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- A Comparative Analysis of Large Language Models for Code Documentation Generation
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Compressing Code Context for LLM-based Issue Resolution
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Automatic Code Documentation Generation Using GPT-3
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- ContextBench: A Benchmark for Context Retrieval in Coding Agents
- MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
- Lost in the Middle: How Language Models Use Long Contexts
- Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
- Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain
- RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository
- Large Language Models Can Be Easily Distracted by Irrelevant Context
- Automatic Code Summarization via ChatGPT: How Far Are We?
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties