Checking Leakage Witnesses versus Certifying Bounded Non-Leakage
cs.CR, cs.CL
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- MAJORITY-3SAT (and Related Problems) in Polynomial Time
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Parameterized Hardness of Zonotope Containment and Neural Network Verification
- The Complexity of Verifying Loop-Free Programs as Differentially Private
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- The Expressive Power of Transformers with Chain of Thought
- The Head Complexity of Boolean Functions in Single-Layer Attention
- The Counting Power of Transformers
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs