2026-09-06-Sun · GimletLabs · InferenceCloud

From Issue 35 (2026-09-06) · 7 stories in this issue

❯ Gimlet Labs Raises $300 Million to Expand Its Multi-Silicon Inference Cloud

Strategic Investors JoinGimlet Labs announced a $300 million Series B led by a16z at a $3 billion valuation. It breaks model inference into tasks that run cooperatively on different types of chips. Total funding is $392 million, with Sapphire Ventures, Menlo Ventures, Arm and Microsoft’s M12 among the participants.

A Larger Round Within MonthsOn March 23, Gimlet announced an $80 million Series A led by Menlo Ventures. By early September, its new round was 3.75 times that size. At the Series A, the company said its customer count had tripled in the five months since its public launch, including a frontier model lab and a hyperscaler, neither named. The rapid return to fundraising comes as agents make sequential model and tool calls, accumulating delays at every step. Customers consequently demand faster responses from inference services. The previous round supported demand validation; the new one must expand delivery at scale.

Different Chips, Different JobsThe company’s product description goes beyond “inference acceleration software.” Its product decomposes model tasks and schedules them across GPUs, CPUs and other accelerators according to customer workload requirements and available hardware. Reading context and generating an answer place different demands on computation and memory, favoring different chips. Gimlet builds cross-chip orchestration software and needs the corresponding data-center connectivity, making it more than a lightweight software tool. Its September announcement says it added billions of dollars in contracted revenue since March and accumulated gigawatts of data-center pipeline. Those are contracts and planned projects, not recognized revenue and live capacity.

Efficiency After ComplexityAssigning tasks to better-suited chips could allow more requests to be processed within the same power constraint, giving an inference cloud an architectural source of profit. Mixed hardware also adds networking, integration, operating and utilization costs. Investors should be buying a whole-system cost advantage, rather than a faster result on one stage of a benchmark. Arm and M12 bring industry connections, but customer bills provide the ultimate test: for the same model and service quality, is the cost per useful task lower, and can that advantage survive the next generation of chips?

▪ SIGNALInference-cloud competition is moving deeper into chip specialization. Performance gains become sustainable margins only after paying for system complexity.