AgentPProf: Semantic Profiler for Long Horizon AI Agents
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/eunomia-bpf/agentsight
Project page: https://arize-ai.github.io/openinference/spec
License: http://creativecommons.org/licenses/by/4.0/
The gist: AI agents increasingly orchestrate long-running activities with users, tools, and system resources for days and weeks.
Terminology
Abstract
AI agents increasingly orchestrate long-running activities with users, tools, and system resources for days and weeks. To improve agent quality, safety, and cost efficiency, developers need to determine where failures happen, what triggers unsafe effects, and which tasks consume the most budget, then optimize those tasks. In systems software, profiling answers similar questions by aggregating resource consumption and attributing it to responsible code paths to identify hotspots. Yet existing agent observability tools focus on per-execution debugging and tracing rather than cross-run, long term profiling, making these questions difficult to answer at scale. Agent observability needs profiling, not only debugging, but profiling agents is challenging: the responsible entities are task intent like diagnose authentication, compare branches rather than code paths, and lack stable identifiers for aggregation. We propose a semantic operation stack model that adapts profiling to agent trajectories. Uniform operations represent all activities, and operation stacks replace the runtime call stack, enabling hierarchical attribution at different granularities. We observe that an agent's task occupies a contiguous span and decomposes into subtasks, so we introduce recursive operation segmentation, which recursively splits trajectories at task boundaries. AgentPProf is a profiler that aggregates agent trajectories into pprof-compatible profiles, enabling flame graph visualization and analysis. AgentPProf reaches 0.764 B cubed F1 against human annotations on CodeTraceBench. On three problem-localization benchmarks, the profile raises MAP by up to 56%, demonstrating that it effectively attributes resources, locates problems, and helps optimize token cost at practical profiling cost. AgentPProf is available at https://github.com/eunomia-bpf/agentsight.
Sources
- AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
- Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems
- How to Interpret Agent Behavior
- Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation
- CodeTracer: Towards Traceable Agent States
- TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
- AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
- TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories
- TraceView: Interactive Visualization of Agentic Program Repair Trajectories
- What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents
- Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories
- HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark
- From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Hodoscope: Unsupervised Monitoring for AI Misbehaviors
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection