2026-08-27-Thu · DeepSeek

From Issue 26 (2026-08-27) · 15 stories in this issue

❯ Zhipu Open-Sources GLM-5.3-Flash; All Public Beta Traffic Ran on Domestic Chip Clusters

IDENTITYOn August 26, Zhipu open-sourced GLM-5.3-Flash and confirmed it was the mystery model Ox Alpha that had been running anonymously on OpenRouter for a week. With 320B total parameters and 18B active, it is MIT-licensed, has a 1 million-token context window, and is the first natively multimodal model in the GLM-5 series. Independent evaluator Artificial Analysis gives it an intelligence index of 57, on par with Claude Opus 4.8, at roughly $0.045 per task.

HARDWAREThe company says all traffic during the anonymous public beta ran entirely on domestic chip clusters. The team built its own inference engine on top of SGLang, tripling end-to-end performance on the same hardware to reach “hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs.” The groundwork was laid earlier: in July, Zhipu acquired Zhongke Jiahe, which traces its roots to the Institute of Computing Technology at the Chinese Academy of Sciences. Its core business is heterogeneous computing software and inference engines, with the goal of raising the utilization of AI chips from multiple vendors. The company did not name the chipmaker.

ARCHITECTURE & PRICINGFor the first time in the GLM main series, the model uses a hybrid of linear attention and sparse attention, combined with IndexPool compression — cutting attention computation to one-third and the key-value cache to about one-quarter. It was pretrained on roughly 30 trillion tokens of multimodal data. On the DeepSWE v1.1 coding benchmark, it scores 63.4, up from 46.2 for the previous generation. API pricing is one-tenth of GLM-5.3’s, and during the limited-time discount period, that is equivalent to one-fortieth of Opus 4.8’s price.

DISTRIBUTIONDuring the anonymous testing window, the model briefly became the most-called model on OpenRouter, more than doubling DeepSeek’s volume within a week — trading free quota for real distributed traffic marks the first time a domestic model has pulled off this playbook. What needs to be recalculated is the model-selection cost for overseas developers: an open-weights option with coding capability on par with top closed-source models, at a forty-fold price difference, is sitting right there — closed-source vendors’ pricing room will be squeezed directly.

▪ SIGNALThe watershed this time isn’t benchmark scores — it’s that a frontier-scale model has, for the first time, proven it can survive without NVIDIA’s supply rhythm.