SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment
cs.LG, cs.AI, cs.CR
Submitted: 2026-09-02
Updated: 2026-09-02
Code: https://github.com/tatsu-lab/stanford_
Terminology
Sources
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- Mixtral of Experts
- Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
- Kimi K2: Open Agentic Intelligence
- Routing-Aware Safety Alignment for Mixture-of-Experts Models
- PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- NVIDIA Nemotron 3: Efficient and Open Intelligence
- gpt-oss-120b & gpt-oss-20b Model Card
- Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent
- Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
- Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought
- BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
- Qwen3 Technical Report
- Towards a Comprehensive Scaling Law of Mixture-of-Experts
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks