Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail
cs.SE, cs.AI
Submitted: 2026-08-24
Updated: 2026-08-24
Comments: 25 pages, 6 figures, 13 tables. Code, data, probe items and rollout logs: https://github.com/Esmail-ibraheem/feedback-that-backfires
Code: https://github.com/Esmail-ibraheem/feedback-that-backfires
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent harnesses record a failed tool call and its error message in the transcript and ask the model to continue, on the assumption that the error is corrective information.
Terminology
Abstract
Agent harnesses record a failed tool call and its error message in the transcript and ask the model to continue, on the assumption that the error is corrective information. We measure whether it is. Defining the corrective gain of a failure record as the change in log-probability of re-emitting the action that just failed, we find the gain is negative for every instruction-tuned model we tested (6 checkpoints, 135M-1.7B, 4 families) in two environments: simulated tool calling and MBPP program repair. Normalised by action length the effect is about-1.03 nats per action token, a factor of 2.8 in the odds of each token, and holds on 90%-100% of individual items, not only on average. Over a fixed candidate set the probability of repeating the failed call rises from 0.06 to 0.54, and greedy decoding reproduces it token for token on 19% of items after the failure versus 0% before. Counterfactuals pairing the same call with a failure message, a success message, or a neutral acknowledgement separate two effects: the failed call's surface form accounts for 83% of the damage, while the semantic contribution of marking it failed is small and inconsistent in sign across environments. The problem is in the harness, not the model's grasp of error messages, and that predicts which remedies work. Replacing the verbatim call with a runtime-generated description of the failure removes 76% of the inversion at no token cost, and making previously-failed strings unreachable at the decoder acts on the same term. Two plausible remedies do not: an explicit "do not repeat" instruction leaves the measured quantity where it was, and deleting the failed attempt to retry from a clean context, the standard prescription for context contamination, is the worst harness we measured for repetition, because it restores the context that produced the failure. The study runs end to end on a CPU; all artefacts are released.
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Program Synthesis with Large Language Models
- Small Language Models are the Future of Agentic AI
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Suppressing Pink Elephants with Direct Principle Feedback
- Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
- Teaching Large Language Models to Self-Debug
- Lost in the Middle: How Language Models Use Long Contexts
- Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
- AgentBench: Evaluating LLMs as Agents
- The Llama 3 Herd of Models
- Self-Refine: Iterative Refinement with Self-Feedback
- The Curious Case of Neural Text Degeneration
- Copy Suppression: Comprehensively Understanding an Attention Head
- When Agents Do Not Stop: Uncovering Infinite Agentic Loops in LLM Agents
- GAIA: a benchmark for General AI Assistants
- Large Language Models Cannot Self-Correct Reasoning Yet
- In-context Learning and Induction Heads
- Do not think about pink elephant!
- Gorilla: Large Language Model Connected with Massive APIs
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties