Hidden State Poisoning Attacks against Mamba-based Language Models
cs.CL, cs.AI, cs.LG
Submitted: 2026-01-05
Updated: 2026-09-01
Comments: 27 pages, 4 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
Code: https://github.com/TortueSagace/hispa
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- BlackMamba: Mixture of Experts for State-Space Models
- Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
- NVIDIA Nemotron 3: Efficient and Open Intelligence
- Defending Against Prompt Injection With a Few DefensiveTokens
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
- Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
- Hymba: A Hybrid-head Architecture for Small Language Models
- Investigating the Indirect Object Identification circuit in Mamba
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Zamba: A Compact 7B SSM Hybrid Model
- Efficiently Modeling Long Sequences with Structured State Spaces
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Prompt Injection Attacks on LLM Generated Reviews of Scientific Publications
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
- Jamba: A Hybrid Transformer-Mamba Language Model
- Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
- Carbon Emissions and Large Neural Network Training
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
- Recurrent Neural Networks (RNNs): A gentle Introduction and Overview
- Locating and Editing Factual Associations in Mamba
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering