WorkBench Revisited: Workplace Agents Two Years On
| Source: arXiv AI
Tags: WorkBench, Claude Opus 4.8, GPT-4, AI agents, workplace automation, AI safety, benchmark
A two-year follow-up on the WorkBench workplace agent benchmark finds Claude Opus 4.8 completing 89% of tasks with harmful action rates of just 2.5% — versus GPT-4's 43% completion and 26% harmful action rate in March 2024 — showing both capability and safety improved together, not in trade-off.
Details
WorkBench tests AI agents on real workplace tasks: sending emails, managing calendars, processing files. In March 2024, GPT-4 was the best agent, completing 43% of tasks and taking an unintended harmful action — such as emailing the wrong person — on 26% of attempts. Two years later, Claude Opus 4.8 completes 89% of tasks and causes harmful actions on only 2.5%. Three findings stand out from the 2026 update. First, capability and safety improved together: the models that complete the most tasks also cause the least unintended harm. There is no observed trade-off. Second, despite the overall progress, frontier models still occasionally make basic irreversible mistakes — sending an email to the wrong recipient still happens. Third, open-weight models have dramatically reduced the cost of achieving the performance level that only proprietary frontier models could reach in 2024. The updated benchmark includes improved data and code quality, new model scores, and an analysis of agent progress across 2024-2026. For enterprise teams deploying workplace AI agents: the 2.5% harmful action rate is still non-trivial for high-stakes workflows, and the finding that irreversible errors persist even in the best model argues for human-in-the-loop checkpoints on actions like sending external emails.