❯ Nvidia Groq 3 LPX Racks Enter Full Production, Nebius First AI Cloud to Adopt
PRODUCTIONNvidia announced on August 24 that its inference accelerator, the Groq 3 LPX rack, has entered full mass production and will be deployed alongside the Vera CPU and Rubin GPU, going live in the data centers of next-gen AI cloud Nebius within the year. These are the first rack products to actually ship since Nvidia acquired Groq for $20 billion.
PERFORMANCEPer the official announcement, a single rack packs 256 LPUs, with 128 GB of on-chip SRAM and 640 TB/s of scale-up bandwidth, built on the MGX architecture with full liquid cooling. In tests by Artificial Analysis using the open-source agentic model Gemma 4 31B, the 100K-token long-context scenario delivers 3,400 tokens/second of output — 4x the closest alternative platform. The previous one-card-handles-all approach is also being split: the Rubin GPU handles context prefill, while the LPU owns the latency-sensitive decode stage, with the two dividing the work inside the same rack.
AGENT LEDGERAlso in the same announcement: Vera Rubin NVL72 results under real agent workloads. Nvidia says that on SemiAnalysis’ AgentX workload running DeepSeek V4 Pro, throughput per megawatt is up to 30x the previous-gen GB300 NVL72, and per-token cost drops by up to 35x. These numbers are aimed squarely at Jalapeño’s efficiency chart — how many concurrent agents you can sustain under a fixed power budget is replacing peak compute as the No. 1 criterion in data center tenders.
▪ SIGNALBy splitting prefill and decode across two chip types, Nvidia concedes that the era of a single GPU handling the entire inference pipeline is over.