Home / News / Z.ai's GLM-5.3 Posts a 50% Coding Jump, Tops Open-Weight Benchmarks
GLM-5.3

Z.ai's GLM-5.3 Posts a 50% Coding Jump, Tops Open-Weight Benchmarks

Aug 14, 20264 min read
Z.ai's GLM-5.3 Posts a 50% Coding Jump, Tops Open-Weight Benchmarks

News Summary

Beijing-based AI lab Z.ai (Zhipu AI) officially launched GLM-5.3 on August 14, 2026 (Beijing Time), positioning the model as its strongest open-weight coding and agentic system to date. The release keeps the same base architecture as GLM-5.2, with every capability gain coming from a much larger post-training phase rather than a bigger foundation model, an approach Z.ai frames as "post-training scaling" pushing the model's intelligence ceiling higher without retraining from scratch.

Availability and Rollout

GLM-5.3 is live immediately through Z.ai's API and the GLM Coding Plan, and has already been rolled out to all existing coding plan subscribers as of the August 14, 2026 (Beijing Time) announcement. The open model weights are not yet public. Z.ai says it plans to release them roughly two weeks later, around the end of August 2026 (Beijing Time), after completing additional safety evaluation and hardening. On the API, GLM-5.3 introduces three selectable thinking-effort levels — low, high, and max — and, in a change from prior releases, no longer allows developers to fully disable the model's reasoning step, which Z.ai flags as a breaking change for applications built around instant, non-reasoning responses.

Coding and Agentic Benchmarks

Z.ai reports GLM-5.3 delivers roughly a 50 percent jump in coding capability over GLM-5.2 on its internal Z.ai Code Bench, moving the model to the top spot among open-weight systems on several public evaluations. Reported benchmark gains include Terminal-Bench 3.0 rising from 4.6 to 28.3, DeepSWE v1.1 climbing from 46.2 to 66.9, and Agents' Last Exam CLI improving from 23.8 to 28.5. On the internal Code Bench, GLM-5.3 scored 31.4 percent using roughly 50,000 tokens of reasoning, ahead of Claude Opus's 29.5 percent at 120,000 tokens, though it trails frontier proprietary systems on the same benchmark. Z.ai also compared the release against GPT-5.6 Sol, DeepSeek-V4 Pro, and Moonshot's Kimi K3 as part of its competitive positioning. The technical gains are attributed to a combination of long-context handling (dubbed IndexShare), a reinforcement-learning technique the company calls SAO, and its "slime" training framework, layered on top of a much larger and more diverse set of long-horizon training environments modeled on professional engineering workflows.

Cybersecurity Capability Grew Faster Than Expected

One notable finding in Z.ai's release notes is that GLM-5.3's offensive-security capability scaled faster than the coding gains that were the primary target of training. On CyberGym, the model scored 84.5 percent, up from GLM-5.2's 77.2 percent, and on ExploitBench it more than doubled its score, from 24.4 percent to 54.4 percent. In timed exploit-development testing, GLM-5.3 completed 105 tasks within two hours and 130 within six hours, compared with 29 and 39 respectively for its predecessor. Z.ai says it has used the model internally to find 2,436 vulnerabilities across 269 open-source projects since GLM-5.2's release, of which 1,097 were rated critical or high severity. At launch, 53 of those findings had been assigned public CVE identifiers, while 2,383 remain under a responsible-disclosure embargo while maintainers are notified. Z.ai ties the delayed release of model weights directly to this finding, stating that additional safety hardening is needed before the underlying weights — which could otherwise be used to replicate the model's offensive capabilities outside of Z.ai's monitored API — are made public.

Industry Context

The release lands amid intensifying competition among Chinese AI developers to close the gap with leading Western labs on coding and agentic tasks, following a string of rapid iterations from Zhipu and rivals over the preceding months. Domestic coverage of the launch, including reports from outlets such as Sina Technology and Eastmoney, framed GLM-5.3 as a significant step for open-source coding models, noting that its coding and agent performance approaches that of top proprietary systems even as it continues to trail the very latest frontier releases from Anthropic and OpenAI. Community discussion ahead of the launch had also centered on requests for native multimodal and vision support, a capability that GLM-5.3 does not introduce in this release.

GLM-5.3Z.ai