PRISM: Recovering Instruction Sets from Language Model Activations
cs.AI, cs.LG
Submitted: 2026-06-08
Updated: 2026-09-27
Code: https://github.com/f/prompts.chat
Terminology
Sources
- Obfuscated Activations Bypass LLM Latent-Space Defenses
- Probing Classifiers: Promises, Shortcomings, and Advances
- Discovering Latent Knowledge in Language Models Without Supervision
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
- Designing and Interpreting Probes with Control Tasks
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- Large Language Models Often Know When They Are Being Evaluated
- LatentQA: Teaching LLMs to Decode Activations Into Natural Language
- Ignore Previous Prompt: Attack Techniques For Language Models
- Generalizing Verifiable Instruction Following
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize
- Instruction-Following Evaluation for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection