Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing
cs.CR, cs.AI, cs.CL, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- ProcessBench: Identifying Process Errors in Mathematical Reasoning
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs