Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
| Source: Hugging Face Blog
Tags: safety alignment, LlamaGuard, refusal calibration, Multiverse Computing, LLM safety, self-distillation
Multiverse Computing's new paper proposes training LLMs to refuse only harmful subsets within a topic — blocking political manipulation while still answering election facts — rather than the current approach of classifying entire topic categories as unsafe.
Details
Current LLM safety systems like LlamaGuard-3 classify harm at the topic level: a prompt touching weapons, elections, or self-harm is treated as uniformly risky. This creates two simultaneous failures: blocking legitimate queries (a civics teacher asking about electoral systems) and missing harmful ones that do not match the blocked keyword pattern (targeted political manipulation framed as a neutral question). Multiverse Computing's paper, 'Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal,' formalizes the problem as a 'narrow-boundary' challenge. The goal is to refuse only a specific harmful subset within a topic — persuasion and manipulation — while continuing to answer the benign complement. The paper identifies a key training difficulty: standard cross-entropy loss that raises refusal on harmful prompts also pushes refusal onto nearby benign ones, causing 'refusal spill.' Their proposed fix is boundary-aware self-distillation — a training approach designed to preserve model behavior in the benign complement while shaping behavior inside the target-harmful subset. Benchmarks XSTest and OR-Bench already measure this over-refusal problem; the paper offers a training-time solution rather than post-hoc recalibration. The work has direct enterprise relevance: the same base model deployed for a general assistant, an educational product, and a public-sector service needs different safety policies within identical topic areas. Per-deployment safety boundaries — not one-size-fits-all topic blocks — are what production systems actually require.