2026-09-25-Fri · Xiaomi · MiMo

From Issue 54 (2026-09-25) · 15 stories in this issue

06 MODEL

❯ Xiaomi unveils MiMo-V3’s core architecture, claiming about one-fifth as much compute to read one million tokens

Xiaomi MiMo lead Luo Fuli introduced HySparse 2, the core architecture planned for MiMo-V3, in an X post. The team says that, compared with MiMo-V2.6’s architecture, it reduces the compute needed to read a one-million-token input to about one-fifth (1/5.02) of the previous amount and reduces the KV cache used while generating an answer by 4.5 times, while improving long-context retrieval benchmarks. These are team-reported tests; MiMo-V3 itself has not launched.

The architecture targets long conversations and agent tasks. Agents accumulate files, dialogue and tool results, then may need to read that history repeatedly to continue working. The computation needed to ingest the material is called prefill; intermediate information retained during answer generation sits in the KV cache. These affect waiting time and memory usage, helping explain why million-token tasks become expensive.

HySparse 2 attempts to select and reuse information more precisely. The team describes moving from block-level to token-level selection and changing how recent context is retained so local and global information can share a cache. It also uses KV bridging and reuse. In practical terms, it aims to reduce repeated computation without losing information needed for the task. The reported benchmark gains do not establish equivalent savings on every user’s bill.

If the improvements hold in real workloads, long-document analysis and agents that repeatedly call tools could use less memory and finish sooner. Saving resources must not come at the expense of missing crucial details, however. Task accuracy, elapsed time and total cost need to be assessed together to establish whether users benefit.

▪ SIGNALLong-context competition now includes affordability and reliability: storing more information matters only if a model can find the useful parts at a reasonable cost.