SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities
cs.CR, cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
- Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
- A Rosetta Stone for AI Benchmarks
- MiniMax Sparse Attention
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
- How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
- The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark
- Kimi K3: Open Frontier Intelligence
- SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories
- Simple synthetic data reduces sycophancy in large language models
- How Predictable Are Large Language Model Capabilities? A Case Study on BIG-bench
- Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis
- SWE-Explore: Benchmarking How Coding Agents Explore Repositories
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs