2026-09-05-Sat · ZAI · DomesticCompute

From Issue 34 (2026-09-05) · 12 stories in this issue

❯ Z.ai’s interim report: inference on 100,000 domestic chips, unit token cost down 80% since January

the disclosureIn its first interim report since listing, Z.ai disclosed large-scale inference on 100,000-class domestic chips and a unit token inference cost down 80% since the start of the year — the first time domestic compute self-sufficiency has appeared in financial-report language. Per SemiAnalysis, 19 of the 60 pages read like a technical blog, with a nine-page glossary attached.

the detailsThe report says a GLM-5.3-powered internal infrastructure agent halved the time needed to optimize inference; the open-sourced GLM-5.3-Flash ran all of its test-period traffic on a domestic chip cluster reportedly sourced from Huawei, Hygon and Moore Threads. That model introduces a hybrid sparse-plus-linear attention architecture that cuts attention compute about 3x versus GLM-5.3 and shrinks the KV cache 4.4x. A week earlier Z.ai had reported fivefold revenue growth against a share price down 60% from its high.

the margin variableAn 80% cost reduction answers the question that dogged last week’s results: whether the API business can produce gross margin. Revenue scaling with calls while cost scales with compute is the affliction of pure-API companies; if inference cost really fell to a fifth of January’s, the margin curve has room to rise. The ones re-weighing are Hong Kong investors pricing Z.ai — the next report’s gross margin will show whether that 80% is an accounting frame or real savings.

▪ SIGNAL100,000 domestic chips plus an 80% cost cut — what Z.ai answered in its report is not a technical question but whether its API business can make money.