Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels
Roman Prosvirnin, Victor Minchenkov, Alexey Soldatov, Vladimir Bashun, Anton Sergeev
cs.CR, cs.AI
Submitted: 2026-08-24
Updated: 2026-08-25
Comments: 17 pages, 12 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Refusal in Language Models Is Mediated by a Single Direction
- Qwen3Guard Technical Report
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
- Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
- A StrongREJECT for Empty Jailbreaks
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
- QLoRA: Efficient Finetuning of Quantized LLMs
- Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
- Quantized Delta Weight Is Safety Keeper
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs