2026-08-07-Fri

From Issue 7 (2026-08-07) · 12 stories in this issue

❯ LatePost Reports ByteDance Discussing a Model with Over 5 Trillion Parameters, the Largest Known in China

SCALE JUMPLatePost reports ByteDance is discussing training a model with more than 5 trillion parameters, surpassing Alibaba’s Qwen3.8-Max at 2.4 trillion and Moonshot AI’s K3 at 2.8 trillion — the largest known in China to date. The plan is still early-stage and may ultimately not be released.

PEOPLE & PATHThe project is led by Seed Foundation head Xiang Liang, in collaboration with LLM pre-training data lead Shen Ke. Xiang Liang joined ByteDance in 2016 and worked on the AML machine-learning middle-platform team before becoming head of the Doubao LLM Foundation team; Shen Ke joined right after graduating from Tsinghua in 2018 and now focuses mainly on pre-training data. The report also cites two directional principles from Zhang Yiming: no distillation, and don’t be swayed by short-term hotspots like coding. In the first half of the year, multiple Seed teams repeatedly reassessed their work; they are now reworking organization and resource allocation — rather than continuing to chase at existing model sizes, they want to push parameters to several times peers’ scale in one move and go straight for the lead.

VERDICTThis is ByteDance’s classic “brute force creates miracles” play, but the cost is out in the open: training and inference costs for a 5-trillion-parameter model will both climb, and HBM and advanced packaging capacity are both in a tight cycle right now. Other domestic model makers need to decide in advance whether to join this round of the parameter arms race — sit out and risk falling behind on capability, or join and bet an entire budget cycle at the most expensive moment for compute. The project isn’t finalized yet, but the pressure has already spread.

▪ SIGNALSkip the catch-up and go straight for scale — ByteDance is betting that compute can buy time.