Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes?
cs.CR, cs.AI, cs.CE, cs.SE
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/open-telemetry/semantic-conventions-genai
Terminology
Sources
- Black-Box Access is Insufficient for Rigorous AI Audits
- Why Do Multi-Agent LLM Systems Fail?
- Visibility into AI Agents
- Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems
- TRAIL: Trace Reasoning and Agentic Issue Localization
- Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
- Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
- Auditable Agents
- Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
- Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures
- DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency
- AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
- Architectural Design Decisions in AI Agent Harnesses
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
- AgentSight: System-Level Observability for AI Agents Using eBPF
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs