An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS
| Source: arXiv AI
Tags: LLM safety, LoRA, AI alignment, harm mitigation, KV cache, IBM Research
Activated LoRA (aLoRA) adapters can be triggered mid-sequence to correct toxic or biased LLM outputs without invalidating the KV cache, providing a modular low-latency safety layer that avoids retraining the base model.
Details
Standard LLM safety approaches are tightly coupled to the model: RLHF and DPO fine-tune the entire model for safety, which is expensive, creates a fixed coupling between safety behavior and model version, and makes it hard to update policies independently. This paper proposes a modular alternative using Activated LoRA (aLoRA) adapters. The framework trains separate expert adapters, each targeting a specific harm type (bias, toxicity, etc.). A learned router inspects intermediate model outputs and dynamically selects the appropriate expert adapter to activate mid-sequence — before the generation continues — without requiring the KV cache to be rebuilt. This keeps latency low while enabling targeted correction rather than blanket safety filtering. The system improves performance on standard safety benchmarks while preserving task performance on unrelated capabilities. The modular design means policies can be updated by swapping or retraining individual expert adapters without touching the base model. The authors are from IBM Research, which has institutional motivation to build enterprise-deployable safety layers. The approach is practically appealing for organizations that need to comply with different safety policies in different deployment contexts without maintaining multiple full model versions.