PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation
cs.CL, cs.AI, cs.SE
Submitted: 2026-09-02
Updated: 2026-09-02
Terminology
Sources
- Program Synthesis with Large Language Models
- CodePlan: Repository-level Coding using LLMs and Planning
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion
- HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- DeepCode: Open Agentic Coding
- AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers
- RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
- Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- Agentless: Demystifying LLM-based Software Engineering Agents
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation
- AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering