The Tokens Remember: When Tokenization Bypasses Knowledge Editing and Unlearning
cs.CR, cs.AI
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- GLM-5: from Vibe Coding to Agentic Engineering
- Olmo 3
- Shieldstral
- Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
- Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
- The Llama 3 Herd of Models
- Qwen3 Technical Report
- Ministral 3
- Gemma 4 Technical Report
- GROM: Gradient-Free Rapid One-Shot Machine Unlearning
- Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation on Digital Ecosystems
- Reverse-Engineering Model Editing on Language Models
- From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs
- FLEKE: Federated Locate-then-Edit Knowledge Editing
- AnyEdit: Edit Any Knowledge Encoded in Language Models
- Lifelong Knowledge Editing requires Better Regularization
- GenRecEdit: Adapting Model Editing for Generative Recommendation with Cold-Start Items
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs