AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents
| Source: arXiv AI
Tags: LLM security, backdoor attacks, model quantization, LLM agents, EMNLP 2026, open-weight models, agentic security
AGENTQ demonstrates that open-weight LLM agents can carry hidden backdoors activating only after standard quantization—reaching up to 100% attack success rate on NF4/FP4/INT8 models while passing pre-deployment audits on the full-precision checkpoint. Accepted at EMNLP 2026.
Details
Quantization is the default deployment path for most open-weight models, but this EMNLP 2026 paper reveals it creates a critical audit gap: an adversary can release a full-precision checkpoint that behaves normally under safety evaluation, yet activates malicious structured function calls once quantized. Unlike text-generation backdoors where harm is mediated by a human reader, agentic backdoors execute autonomously without oversight.\n\nNaive adaptation of prior backdoor methods produces post-quantization malicious behavior but degrades benign utility—making attacks detectable. AGENTQ solves this by combining layer-banded LoRA injection with partial-PGD repair over a multi-codebook quantization-equivalence class. This concentrates malicious behavior in quantization-sensitive layers while preserving normal agentic capability.\n\nAcross three trigger-action pairs and three quantization codebooks (NF4, FP4, INT8), AGENTQ reaches up to 100% post-quantization attack success rate with minimal benign utility loss. The authors argue this makes quantization-aware safety evaluation a necessary requirement before deploying open-weight agents—a standard that does not currently exist in the field.