Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks

| Source: arXiv AI

Tags: LLM safety, jailbreak defense, open-weight models, representation engineering, alignment

Bait-and-Recover defeats white-box representation engineering attacks on open-weight LLMs by poisoning the observation path: a bait adapter corrupts the signal attackers measure, a recovery adapter restores clean computation, raising minimum refusal rate from 16.25% to 71.75% with negligible capability loss.

Details

Open-weight LLMs face a specific class of attack: representation engineering jailbreaks that estimate refusal directions and apply projection-matrix edits to suppress safety alignment, executable in minutes on a single GPU without gradient training. This paper proposes Bait-and-Recover, a weight-level defense. The mechanism uses two paired adapters: a bait adapter placed where attackers typically read activations, and a recovery adapter at the subsequent layer. Training via gradient routing decouples the observation path from the behavior path. By actively poisoning the residual stream used for measurement, the defense disrupts the attacker's edit search. The recovery adapter then restores clean downstream computation. Evaluated across four open-weight models under a strict behavior-preservation budget (KL ≤ 0.10), Bait-and-Recover raises the minimum refusal rate against white-box edit searches from 16.25% to 71.75% with negligible impact on general benchmarks. Code is publicly released. This is particularly relevant as open-weight models (Llama, Mistral, Qwen, etc.) proliferate in enterprise deployments. The low-cost attack threat (single GPU, minutes) means the defense asymmetry was strongly in the attacker's favor before this work.