❯ Zhipu GLM-5.3 Goes Live on API, Independent Score of 60 Ties Kimi K3
LAUNCH & PRICINGZhipu’s GLM-5.3 is now officially available on the official API and partner gateways, with pricing unchanged from GLM-5.2, targeting coding, defensive cybersecurity, and long-horizon agent tasks. This generation keeps the same base model — all gains come from post-training scaling: longer training runs, a training environment dozens of times larger than the previous generation, and a broader mix of environment types. Open-source weights won’t be released until security evaluation and hardening are complete.
BENCHMARKSIndependent evaluator Artificial Analysis gives it an Intelligence Index of 60, tying Kimi K3 and up 7 points from the previous GLM-5.2, though it still trails Opus 5’s 63 and Fable 5’s 62. Once the weights are opened, it will be tied for first among open-source models. According to Zhipu’s own published figures, Terminal-Bench 3.0 jumped from 4.6 to 28.3, DeepSWE rose from 46.2 to 66.9, and the CyberGym vulnerability discovery rate hit 84.5%.
DISSENTResearcher teortaxesTex takes a different view, arguing this generation is overly skewed toward software engineering, with a slight regression on CritPt; overall, Kimi K3 remains the most well-rounded among Chinese models. This divergence is worth watching: scores built up through post-training in specific environments may not hold up when transferred to other tasks. The same 60 points don’t carry equal weight whether they sit on a coding pipeline or on research reasoning — before picking a model, it’s worth first measuring how much your own tasks overlap with the evaluation environment.
▪ SIGNALGaining 7 points on the same base with post-training alone shows the decisive lever in this round of competition has moved from pretraining scale to the ability to construct training environments.