The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play
Michael Macaulay, Harmony Bouabid, Guo Gen Ang, Sasha Shaw
cs.AI, cs.CR, cs.CY
Submitted: 2026-07-28
Comments: 20 pages, 3 figures, 1 table
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
- InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection