CryptanalysisBench: Can LLMs do Cryptanalysis?
Lukas Fluri, Avital Shafran, Nicholas Carlini, Matthew Jagielski, Milad Nasr, Orr Dunkelman, Eyal Ronen, Florian Tramèr
cs.CR
Submitted: 2026-07-20
Comments: 46 pages, 5 figures, 4 tables
Code: https://github.com/harbor-framework/harbor
Project page: https://meta-llama.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Cryptanalysis - the task of finding attacks against cryptographic schemes - sits at the intersection of mathematical reasoning and cybersecurity, two areas where LLMs have advanced fastest.
Terminology
Abstract
Cryptanalysis - the task of finding attacks against cryptographic schemes - sits at the intersection of mathematical reasoning and cybersecurity, two areas where LLMs have advanced fastest. Cryptanalysis represents both a clean testbed for frontier reasoning (as practical attacks can be automatically verified) and a domain with unusually high stakes, since the primitives under study underpin our digital security. In this paper we ask whether LLMs can do cryptanalysis, and find that the answer is increasingly yes. We introduce CryptanalysisBench, 191 tasks across six families of cryptographic primitives (block ciphers, hash functions, etc.) drawn primarily from four NIST standardization competitions. Our benchmark consists of three tiers: (i) primitives with known practical breaks; (ii) primitives with no known practical break, evaluated both at full strength and as scaled-down variants; and (iii) a challenge set of production primitives at the frontier of cryptanalysis. Five frontier models (Claude Opus 4.8, Sonnet 5, Mythos 5, GPT 5.5, and the open-weights GLM 5.2) break 65%-86% of Tier 1 schemes, 6-12 Tier-2 schemes at full strength, and 24-61 across all scaled-down variants. Beyond deriving known results, models produce novel cryptanalysis, such as a key-recovery attack that exploits a design flaw in the SpoC AEAD and an error in KINDI's published CCA-security proof, both to the best of our knowledge not previously known. We release CryptanalysisBench as a tool to help track if (or when) AI cryptanalysis becomes a serious factor and as a scaffold for stress-testing candidate schemes before deployment. The attacks that the benchmark already surfaces are an early snapshot of a fast-moving frontier that may soon match, and in places exceed, the published state of the art.
Sources
- SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
- Primitive sets and von Mangoldt chains: Erd\H{o}s Problem #1196 and beyond
- Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
- Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
- CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
- Generalized Power Attacks against Crypto Hardware using Long-Range Deep Learning
- eyeballvul: a future-proof benchmark for vulnerability detection in the wild
- CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering
- CryptoQA: A Large-scale Question-answering Dataset for AI-assisted Cryptography
- LLM Agents can Autonomously Exploit One-day Vulnerabilities
- LLM Agents can Autonomously Hack Websites
- Unsupervised Cipher Cracking Using Discrete GANs
- When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
- Endless Jailbreaks with Bijection Learning
- Can Transformers Break Encryption Schemes via In-Context Learning?
- ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation
- VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs