AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMs
cs.CR
Submitted: 2026-09-01
Updated: 2026-09-03
Terminology
Sources
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- Reasoning Models Don't Always Say What They Think
- Training Verifiers to Solve Math Word Problems
- Inference-Time Backdoors via Chat Templates: From LLM Supply Chains to Agentic System Compromise
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity
- BadEdit: Backdooring large language models by model editing
- Release Strategies and the Social Impacts of Language Models
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs