An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS
cs.LG, cs.AI, cs.CL, cs.CY
Submitted: 2026-09-12
Updated: 2026-09-19
Code: https://github.com/vllm-project/vllm
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs.
Terminology
Abstract
Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. Existing alignment methods, while effective, are costly and tightly coupled to the model, limiting flexibility and scalability. We propose a modular correction framework that augments pretrained LLMs with Activated LoRA (aLoRA) adapters and a context-aware routing mechanism to eliminate harms from misaligned model responses. Our approach enables expert adapters to activate mid-sequence without invalidating the KV cache, allowing low-latency, targeted correction during generation. Each expert is trained to detect and mitigate specific harms, such as bias or toxicity. A learned router dynamically selects appropriate experts based on the models intermediate outputs. We demonstrate that our system improves alignment on standard safety benchmarks while preserving task performance, offering a lightweight and efficient path toward safer and more controllable LLM deployments.
Sources
- Foundational Challenges in Assuring Alignment and Safety of Large Language Models
- Language Models are Few-Shot Learners
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- A Survey on Mixture of Experts in Large Language Models
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Mixture-of-LoRAs: An Efficient Multitask Tuning for Large Language Models
- Activated LoRA: Fine-tuned LLMs for Intrinsics
- LoRA: Low-Rank Adaptation of Large Language Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
- MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
- Aligners: Decoupling LLMs and Alignment
- Training language models to follow instructions with human feedback
- Granite Guardian
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks