❯ Google TPUv7 gets first third-party inference results; SemiAnalysis reports up to 50% better performance per dollar than Nvidia in selected configurations
Published comparisonChip research firm SemiAnalysis published an InferenceX preview for Ironwood on September 7, using Qwen3.5 397B as the initial model. Its headline advantage compares FP8 against FP8 on B200/B300. The result depends on specified serving configurations and cost assumptions; it is not a Google cloud-service price cut. Test report.
Software accessThe more consequential development is TorchTPU: developers can retain PyTorch and familiar inference frameworks when serving open-weight models on TPUs, reducing cross-framework adaptation. According to the research, the software remains in private beta, with open sourcing planned around mid-October. Results for multi-turn agent workloads are still forthcoming. Dylan Patel’s concurrent comments come from the same research team, not a second independent test.
Competitive gapMuch of Nvidia’s advantage comes from software that can run models as soon as they are released. Improving adaptation speed would help TPU operating-cost advantages translate into external orders. The results give cloud providers a concrete alternative, but moving from a handful of models to routine procurement requires broader model coverage. A strong benchmark cannot replace a complete developer toolchain.
▪ SIGNALWinning external inference business requires TPU to answer two questions: how much each dollar buys, and how soon a newly released model can run.