The Deception Delta: Adversarial Evaluation of LLM-Based Smart Contract Bytecode Forensics
cs.CR
Submitted: 2026-09-12
Updated: 2026-10-06
Comments: 14 pages, 1 figure, 4 tables. Accepted at ACM CCS 2026
Project page: https://ethereum.github.io/yellowpaper
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models are increasingly used in blockchain forensic investigations to interpret unverified smart contract bytecode.
Terminology
Abstract
Large language models are increasingly used in blockchain forensic investigations to interpret unverified smart contract bytecode. Their robustness has not been systematically tested against contracts adversarially designed to mislead analysis. We evaluate 22 frontier models on 13 purpose-built contracts (9 deception vectors, 4 controls) across six prompt strategies, yielding 8,528 analyzable non-refusal runs against contracts with EVM-verified ground truth. A calibrated LLM-as-judge pipeline, supported by two judge-independent metrics and 50 human gold-standard labels, shows that adversarial deception reduces drain detection by 20.0 percentage points (95% CI: [17.2, 22.8]) relative to functionally matched controls. Structural camouflage via multi-hop call chains, XOR-masked selectors, and storage-loaded drain parameters resists detection across nearly all models. Beyond non-detection, we identify rationalization: models correctly describe the hidden drain mechanism but accept the contract's deceptive framing and dismiss it as benign, yielding positive but incorrect evidence of safety. Simple guard instructions provide no aggregate benefit and destabilize individual models in both directions. Structural deception is largely insensitive across the six tested prompt strategies, more consistent with a capability limitation than with a simple prompting problem. Only five models from two providers exceed 50% detection. Under our single-shot, raw-bytecode-only protocol, current LLMs are not reliable standalone forensic tools. Our central claim does not extend to multi-turn, tool-augmented, source-aware, or decompiler-in-the-loop workflows; a source-code boundary check is reported as an explicit subset analysis rather than as part of the main evaluation.
Sources
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- Do you still need a manual smart contract audit?
- Decompiling Smart Contracts with a Large Language Model
- Precise Static Identification of Ethereum Storage Variables (Extended Version)
- CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- Revealing Adversarial Smart Contracts through Semantic Interpretation and Uncertainty Estimation
- Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- TraceLLM: Security Diagnosis Through Traces and Smart Contracts in Ethereum
- LLM-SmartAudit: Advanced Smart Contract Vulnerability Detection
- RPHunter: Unveiling Rug Pull Schemes in Crypto Token via Code-and-Transaction Fusion Analysis
- Large Language Model based Smart Contract Auditing with LLMBugScanner
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs