Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks
| Source: MarkTechPost
Tags: GLM-5.3, Z.ai, foundation models, coding AI, cybersecurity AI, post-training, LLM benchmarks
Z.ai released GLM-5.3 on August 14, 2026 — same 743B base model as GLM-5.2, all gains from post-training alone — with Terminal-Bench 3.0 jumping from 4.6 to 28.3 and CyberGym cybersecurity scoring reaching 84.5%, edging past GPT-5.6 Sol and Mythos 5 on that benchmark.
Details
Z.ai's GLM-5.3 demonstrates that post-training scaling alone — without touching the base model — can deliver substantial capability improvements. The release reuses GLM-5.2's 743B parameter weights, with all gains attributed to more task environments, more environment types, and longer training runs. The most striking results are in coding and cybersecurity. Terminal-Bench 3.0 (long-horizon CLI tasks) improved from 4.6 to 28.3 — roughly a 6x jump. DeepSWE v1.1 (software engineering) moved from 46.2 to 66.9. On cybersecurity, CyberGym rose from 77.2% to 84.5%, outperforming both GPT-5.6 Sol and Mythos 5. ExploitBench — full exploitation chain reasoning — jumped from 24.4% to 54.4%. The cybersecurity gains were described by Z.ai as unplanned: adding vulnerability-discovery data for single-bug reasoning caused capabilities to compound unexpectedly, with the model beginning to form coherent multi-step exploitation plans. GLM-5.3 is immediately available via the Z.ai API and GLM Coding Plan. Weights are expected roughly two weeks post-launch after safety evaluation. On independent benchmarks, GLM-5.3 still trails GPT-5.6 Sol and Fable 5 on several harder coding tasks. All figures are vendor-reported with documented harness and sampling settings.