Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark

arXiv:2609.25647 · cs.AI · Submitted 2026-09-22 · Read on arXiv

cs.AI

Submitted: 2026-09-22

Updated: 2026-09-22

Comments: 26 pages, 4 figures, 3 tables

Code: https://github.com/caoyanze426-crypto/early-outcome-calibration-audit

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers