Reading Zhipu’s GLM-5.3 results past the headline number
| Source: AI News (ainews.com)
Tags: GLM-5.3, Zhipu, CyberGym, ExploitBench, cybersecurity, open-weights, Anthropic, OpenAI, benchmarks
Zhipu's GLM-5.3 edges out Anthropic Mythos 5 and OpenAI GPT-5.6 Sol on the CyberGym vulnerability-finding benchmark by 0.7 points, but trails by 24 percentage points on ExploitBench — a gap Zhipu acknowledges in its own release note while the coverage focused only on the headline win.
Details
Zhipu AI launched GLM-5.3 on August 14, a coding-focused model whose cybersecurity results generated headlines claiming a Chinese open-weights model now surpasses American frontier labs at bug hunting. The reality is more layered. On CyberGym — which tests whether a model can identify a vulnerability from readable source code — GLM-5.3 scores 84.5% versus 83.8% for Anthropic's Mythos 5 and 83.6% for OpenAI's GPT-5.6 Sol. That 0.7-point margin is what most coverage amplified. Two other benchmarks in Zhipu's own release document tell a different story. ExploitBench, which requires reasoning about real vulnerabilities and their exploitation paths, puts GLM-5.3 at 54.4% against Mythos 5's 78.0% and GPT-5.6 Sol's 76.5% — a 24-point gap. On ExploitGym, which measures exploitation tasks completed within a fixed time window, GLM-5.3 finishes 105 tasks in two hours and 130 in six; Mythos 5 completes 181 and 247 respectively. Zhipu is notably candid about this in its release note, writing that capability is 'growing fastest exactly where we are furthest behind.' The distinction matters: finding a vulnerability and building a working exploit from it are meaningfully different capabilities. GLM-5.3 weights are planned for public release, while Anthropic's equivalent vulnerability research sits behind restricted access for precisely this reason.