❯ Financial Times: ByteDance Is Pretraining a Model With Up to 10 Trillion Parameters — Zhang Yiming Demands No Distillation
SCALEAccording to the Financial Times, ByteDance is pretraining a model with up to 10 trillion parameters, led by a roughly 2,000-person Seed team. That scale is about 3 times that of Moonshot AI’s Kimi K3, and exceeds outside estimates of Anthropic’s Mythos 5 at roughly 8 trillion parameters.
NO SHORTCUTSThe report says founder Zhang Yiming has explicitly demanded pretraining from scratch — no distillation. The model is still in the pretraining phase, which typically runs three to six months; the final parameter scale is locked in afterward, followed by a decision on fine-tuning and release. For a company long accused of taking shortcuts, training a frontier model from scratch is the most expensive — and hardest to rebut — answer. ByteDance already proved once this year that its self-developed route works in video generation; this time it is bringing the same playbook to general-purpose models.
COMPUTE FIRSTThe parameter race was assumed to have cooled under the efficiency route — now China and the US are both pulling it back at the same time. What tightens first is compute scheduling: a single 10-trillion-parameter pretraining run will consume cluster time equivalent to ByteDance’s next six months of investment in recommendations and video generation — investment that now has to get back in line. Whether the model ever ships is a later question; training resource allocation has already moved.
▪ SIGNALTraining from scratch is, right now, the only move that the accusation of ‘copying’ can’t touch.