Schr"odinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?
cs.SE, cs.AI
Submitted: 2026-08-21
Updated: 2026-08-21
Code: https://github.com/cslsolow/Schrodinger-Repo
Terminology
Sources
- Pruning the Unsurprising: Efficient LLM Reasoning via First-Token Surprisal
- Dockerless: Environment-Free Program Verifier for Coding Agents
- SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents
- Test vs Mutant: Adversarial LLM Agents for Robust Unit Test Generation
- Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution
- Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
- LLM Agents Can See Code Repositories
- SWE-bench Goes Live!
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
- SWE-Exp: Experience-Driven Software Issue Resolution
- SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
- SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
- SWE-Bench+: Enhanced Coding Benchmark for LLMs
- Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties