Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
cs.CR, cs.AI, cs.SE
Submitted: 2026-09-27
Updated: 2026-10-03
Code: https://github.com/yxsec/system-one-security-eval
Terminology
Sources
- Confidently Wrong: Severity-Aware Calibration of Prompt-Injection Detectors under Attack Shift
- Defeating Prompt Injections by Design
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
- Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
- WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
- JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places
- Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It
- ProbGuard: Proactive Runtime Monitoring for LLM Agent Safety via Probabilistic Prediction
- Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions
- REFLEX with Jev for Efficient Selective Control in LLM Agents
- JevOut: Natural Context Can Flip Decision Models
- Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs