The Implications of Linguistic Illegibility for LLM Security
cs.LG, cs.CR
Submitted: 2026-09-02
Updated: 2026-10-06
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
- Constitutional AI: Harmlessness from AI Feedback
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Discovering Latent Knowledge in Language Models Without Supervision
- Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents
- How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
- Securing AI Agents with Information-Flow Control
- Defeating Prompt Injections by Design
- Toy Models of Superposition
- Scaling and evaluating sparse autoencoders
- AI Control: Improving Safety Despite Intentional Subversion
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- Emergent Linear Representations in World Models of Self-Supervised Sequence Models
- Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
- Partitioned Tags, Shared Data: Reconciling Strict Cache Isolation with Write-Shared Coherence
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks