JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges
cs.CR
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs