AuraForge: Scaling Security Supervision for Training Coding Agents
cs.CR, cs.CL, cs.CY
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/github/advisory-database
Terminology
Sources
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
- Vibe Coding in Practice: Motivations, Challenges, and a Future Outlook -- a Grey Literature Review
- ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
- Vibe coding: programming through conversation with artificial intelligence
- CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities
- Detecting Safety Violations Across Many Agent Traces
- BaxBench: Can LLMs Generate Correct and Secure Backends?
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
- PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- SWE-smith: Scaling Data for Software Engineering Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs