Safin-1: Safety from Within through Memory-Native State Evolution
cs.LG
Submitted: 2026-08-31
Updated: 2026-08-31
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions.
Terminology
Abstract
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family of foundation models realizing this principle through memory routing and state evolution. Safin-1 is built on Memory-Anchor Routing across Context History (MARCH), a network architecture that maintains structured memory states and selectively retrieves relevant historical information through content-conditioned routing. It supports test-time adaptation of persistent capability states without repeatedly modifying the backbone, enabling controlled specialization over a shared foundation. We investigate this interface on downstream safety tasks through a Safety State, demonstrating effective state-based adaptation with substantial safety improvements. More broadly, the routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior. Evaluations across general capabilities, long-context understanding, retrieval, and efficiency further validate Safin-1. These findings provide a path toward safety as a state-native and adaptively maintainable capability. This work is only an initial architectural exploration of Safety from Within, and substantial further work is needed to realize this broader vision.
Sources
- \texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World
- Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
- SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law
- DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
- EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
- TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
- MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Linear Transformers Are Secretly Fast Weight Programmers
- Scaling Linear Attention with Sparse State Expansion
- Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- MoM: Linear Sequence Modeling with Mixture-of-Memories
- Dynamic Linear Attention
- MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
- Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
- MoBA: Mixture of Block Attention for Long-Context LLMs
- Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks