ACTS: A multi-tier benchmark evaluating LLM cipher identification under controlled blind conditions
cs.CR
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/ReFirmLabs/binwalk
Terminology
Sources
- Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities
- To Err is Machine: Vulnerability Detection Challenges LLM Reasoning
- PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
- Machine learning the electronic structure of matter across temperatures
- AI-based Dynamic Schedule Calculation in Time Sensitive Networks using GCN-TD3
- Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
- Train 'n Trade: Foundations of Parameter Markets
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Confabulation: The Surprising Value of Large Language Model Hallucinations
- Accelerating Post-Quantum Cryptography via LLM-Driven Hardware-Software Co-Design
- Directed Social Regard: Surfacing Targeted Advocacy, Opposition, Aid, Harms, and Victimization in Online Media
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Anchored Confabulation: Partial Evidence Non-Monotonically Amplifies Confident Hallucination in LLMs
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs