HazardAuditor: From Executable Threats to Safer Computer-Use Agents
| Source: arXiv AI
Tags: agent safety, computer-use agents, guard models, GuardPO, Claude Code, AI safety, LLM agents
HazardAuditor is a new execution-grounded safety framework for computer-use agents (browsers, terminals, file systems) that improves guard model accuracy by up to 16.5 percentage points over prior methods by introducing Guard Policy Optimization to fix structural training mismatches.
Details
Computer-use agents that interact with browsers, terminals, file systems, and external services introduce safety risks that emerge through runtime behavior, not just static text outputs. Existing guard models are designed for static prompt-response pairs and fail in agent execution contexts. HazardAuditor closes two gaps: it runs heterogeneous agents (tested on Claude Code, Codex, Hermes, and OpenClaw) in controlled environments and normalizes their interactions into a canonical event format for cross-framework supervision. The paper also identifies a structural training problem with generative guard models: token-level post-training objectives cause longer rationales to dominate gradient updates, weakening the actual safety verdict signal. The proposed Guard Policy Optimization (GuardPO) converts deterministic safety outcomes into sequence-level advantages and normalizes rationale and verdict regions, making the safety decision the effective optimization target. Across multiple benchmarks and all five tested agent frameworks, HazardAuditor improves accuracy by up to 16.5 percentage points over the strongest prior guard. Code, models, and evaluation artifacts are being released publicly. This is directly relevant to anyone building or auditing agentic systems that interact with the real world.