❯ Qwen3.8-Max Takes Fifth in Intelligence Index, First in Agentic Index, but Costs Still Higher than Kimi K3
SCORESThe latest results from independent evaluation body Artificial Analysis show Alibaba’s Qwen3.8-Max scoring 56 in the Intelligence Index, ranking fifth, and taking first in the Agentic Index; average cost per completed task is USD 1.14. Alibaba’s Tongyi official account confirmed both rankings on X that day, citing the agency’s public leaderboard.
CONTRASTBut the leaderboard has another half: the open-weights camp’s leader Kimi K3 scores 1 point higher, with a per-task cost of just USD 0.86 — roughly 25% lower than Qwen3.8-Max. In other words, at the same level of intelligence, this closed-source version has not bought a cost advantage. Looking at the domestic lineup, Qwen3.8-Max has 2.4 trillion parameters and K3 has 2.8 trillion; the two are so close that they’re separated by a single decimal place.
TAKEAWAYEnterprise buyers choosing models are now really comparing cost per task, not leaderboard rankings. When the score gap is only 1 point but the price gap is 25%, ranking position is no longer a procurement rationale. Taking first in the Agentic Index is a more tangible win for Alibaba — agent scenarios demand high stability in multi-turn tool calls, and being first in this category converts to orders better than a higher overall score.
▪ SIGNALA 1-point score gap and a 25% price gap — no need to hesitate over which way procurement decisions will tilt.