Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
cs.AI, cs.CL, cs.CY, cs.LG
Submitted: 2026-03-08
Updated: 2026-09-30
Code: https://github.com/davidgringras/safety-under-scaffolding
Terminology
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- International AI Safety Report 2026
- Answer Matching Outperforms Multiple Choice for Language Model Evaluation
- Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
- A Survey on LLM-as-a-Judge
- OLMES: A Standard for Language Model Evaluations
- Pre-registration for Predictive Modeling
- Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
- The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
- SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
- Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
- SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
- Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
- Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
- AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
- OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
- Safety in Large Reasoning Models: A Survey
- LLMs May Perform MCQA by Selecting the Least Incorrect Option
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection