Models
GLM-5.3 Ties for the Open-Model Crown; Weights Wait
Z.ai's GLM-5.3 ties Kimi K3 atop the open-model rankings and undercuts it on cost per task, but the open weights are delayed: the model proved unexpectedly good at finding security vulnerabilities.

When was the last time a lab delayed an open-weight release because the model was too good at something? That is the situation Z.ai finds itself in with GLM-5.3, which scored 60 on the Artificial Analysis Intelligence Index, tying Kimi K3 for first place among open models, while its weights sit in a holding pattern for a reason nobody used to cite: the model turned out to be unexpectedly effective at finding security vulnerabilities in software.
The scoreboard
The benchmark story goes beyond the open-model bracket. On GDPval-AA v2, which measures performance on professional work tasks, GLM-5.3 posted a 1,770 Elo rating, second across all models, open or closed, behind only Claude Opus 5 at 1,855. On cost, Artificial Analysis measured $0.68 per task, 19 percent cheaper than Kimi K3's $0.84. Cost per task is the honest metric here: a model with cheap tokens that thinks in long chains can still produce an expensive bill, and this measure captures that. The same measure also shows the direction of travel, since GLM-5.3 costs 54 percent more per task than its predecessor GLM-5.2. Open no longer means cheap by default; capable open models are getting pricier as they do more work per query.
Why hold the weights back?
Z.ai says it will delay the open-weight release by roughly two weeks, harden its safety checks in the meantime, and limit full capability access to selected security partners. The model remains usable through the Z.ai API today. Skeptics might call the delay marketing; the concern it points to is not hypothetical. Earlier this month, an autonomous attack campaign was traced to an agent running on an open-weight model, after closed providers' safety controls had refused the same task. Weights, once published, cannot be recalled. A staged release from a Chinese lab, of all places, suggests the open-weight calculus is shifting industry-wide.
What we take from it
Two things can be true at once. For any organization wanting to run models on its own infrastructure, for cost, privacy, or independence, the open ecosystem now offers near-frontier capability, and GLM-5.3's professional-task showing is the strongest evidence yet; the same week, one frontier tracker concluded Chinese labs have largely closed the gap with Western ones. And the launch-week numbers deserve their usual quarantine: these are vendor-adjacent measurements, and independent replication typically takes a few weeks to settle. If a model this capable at vulnerability hunting ships openly, the copy in a defender's hands is identical to the one in an attacker's; access controls and monitoring around self-hosted models stop being optional hygiene.
Sources: The Decoder

Written by
Muhammet Fatih Batman
Founder & Editor
Founder of YZ Uzman, with 20+ years of experience in web design and software development.