2026-08-25-Tue

From Issue 24 (2026-08-25) · 15 stories in this issue

❯ NVIDIA Groq 3 LPX Enters Mass Production: 100K-Token Long Context Benchmarked at 3,400 Tokens/Sec

PRODUCTION & BENCHPer NVIDIA’s announcement, the dedicated inference accelerator Groq 3 LPX has entered full mass production. In Artificial Analysis’s standard benchmark, running the open-source model Gemma 4 31B with a 100K-token active context, it delivered 3,400 tokens per second output — 4x the next-best platform under the same long-context conditions, and the fastest result ever recorded for this model. A single rack can hold up to 256 LPUs, with 128GB of high-bandwidth SRAM.

BACKSTORYThe chip is positioned as a complement to the Vera Rubin NVL72 platform, targeting agent workloads that need sustained token generation over long stretches. What’s more notable is the timeline: per The Register, eight months have passed since NVIDIA secured a non-exclusive technology license from Groq — the first time Groq technology has appeared in an NVIDIA rack-scale product. Cloud provider Nebius is the first to deploy it, making it available through its Token Factory service.

INFERENCE SPLITTraining and inference are being split into two independent hardware product lines, making tokens per second at long context an independently priced metric. The first constraint to loosen is the response-latency budget for agent products: an agent chaining dozens of sequential tool calls was previously limited by generation speed to asynchronous interaction. At this throughput level, synchronous real-time product forms become viable for the first time — and per-task inference costs are being recalculated accordingly.

▪ SIGNALPer The Register, that roughly $20 billion license has, for the first time, taken a shape that benchmarks can record.