Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development
cs.CL, cs.AI, cs.SE
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 6 pages, accepted to The 2nd Workshop on Explainable and Reliable Software Systems (EXPRESS 2026)
Code: https://github.com/illinoisdata/STALE-bench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Parallel coding agents can produce patches that work alone but fail when merged.
Terminology
Abstract
Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on. We study these failures with stale, a benchmark for semantic coordination. Our evaluation runs the same tests on each patch alone and on their combination, counting only failures introduced by combining the patches. We use three tiers: synthetic tasks with controlled interface changes, pairs of merged pull requests, and constructed tasks that use real Django helpers. Among 834 runs on 417 mined Django pairs, only one showed interference after correcting the grading procedure. On constructed tasks using 12 Django helpers, interference occurred in 97% of runs. A message describing the completed concurrent change recovered 82% of runs. Reviewed pull requests may contain few unresolved parallel changes, even when agents fail on controlled tasks using real code. The constructed failure rates do not estimate how often these problems occur in practice.
Sources
- AgenticFlict: A Large-Scale Dataset of Merge Conflicts in AI Coding Agent Pull Requests on GitHub
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering