❯ Zhipu Launches GLM-5.3-FlashX at Up to 200 Tokens a Second for 2.5 Times the Base Price
Speed TierZhipu said it has opened the GLM-5.3-FlashX API, with peak output of 200 tokens a second, roughly five times GLM-5.3-Flash. The model retains a 1-million-token context window and multimodal input; the main differences are service speed and billing.
Paying for LatencyIn a comparison with the base tier, Zhipu lists FlashX input and output at CNY2 and CNY7 per million tokens, about 2.5 times the base price. Paying 2.5 times for five times the speed may suit real-time support, voice interfaces and high-concurrency agents; offline batch workloads retain a cheaper option. Enterprises must still verify concurrency quotas and peak-period commitments.
Procurement ChoiceDevelopers can now assign a price to latency inside one model family, avoiding a model change and full intelligence reevaluation. Actual gains still depend on time to first token, concurrency and peak throughput; a stated 200-token peak is not the same as stable production speed. Testing with real traffic before routing every request to the faster tier will usually cost less.
▪ SIGNALFlashX is not selling a smarter model but priced waiting time; real-time products must recover the premium through higher completion rates or more concurrency.