trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
cs.CL, cs.AI, cs.SE
Submitted: 2026-08-29
Updated: 2026-10-02
Comments: 16 pages (8-page main text). Under review at a NeurIPS 2026 workshop. Code, data, and raw verdicts: https://github.com/mohammadi-hadi/trajectory-judge
Code: https://github.com/mohammadi-hadi/trajectory-judge
License: http://creativecommons.org/licenses/by/4.0/
The gist: Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final reply and ask whether it was handled well.
Terminology
Abstract
Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final reply and ask whether it was handled well. The metric is structurally blind to an agent that reaches the right answer the wrong way. We measure that blind spot where ground truth is known by construction: a deterministic tool-using support-desk environment, a scripted oracle policy that always solves it, and a fault injector that breaks exactly one thing at a known step, stratifying faults by whether the customer-visible outcome survived (silent) or not (loud). Five judges (programmatic rules, outcome-only, step-rubric at two model sizes, and a self-consistency ensemble) are scored on detection, step localisation, fault typing, calibration, and cost over 400 trajectories. The outcome-only judge catches 84% of loud faults but 45% of silent ones while flagging 33% of correct trajectories; a step-rubric judge reaches 77% silent recall with zero false alarms at 3x the cost. No judge reads the final reply: an invented promise appended to an otherwise perfect trajectory evades the rules entirely and the step judge 82% of the time, and self-consistency triples cost while improving nothing. We argue that judge evaluations must stratify recall by outcome survival, and release the environment, the injector, all raw verdicts, and an analysis pipeline that rebuilds every number offline.
Sources
- Why Do Multi-Agent LLM Systems Fail?
- TRAIL: Trace Reasoning and Agentic Issue Localization
- The Llama 3 Herd of Models
- A Survey on LLM-as-a-Judge
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Language Models (Mostly) Know What They Know
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
- Let's Verify Step by Step
- AgentBench: Evaluating LLMs as Agents
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
- Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference
- EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
- LLM Evaluators Recognize and Favor Their Own Generations
- Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
- Solving math word problems with process- and outcome-based feedback
- Large Language Models are not Fair Evaluators
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering