ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
cs.AI, cs.CR
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/dreadnode/scopebench-pilot
Terminology
Sources
- ReAct: Synergizing Reasoning and Acting in Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
- Refusal in Language Models Is Mediated by a Single Direction
- PentestJudge: Judging Agent Behavior Against Operational Requirements
- ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
- CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection