❯ SemiAnalysis benchmarks OpenAI’s in-house Jalapeño chip ahead of Blackwell in most cases
the numbersSemiAnalysis’s verdict on OpenAI’s in-house accelerator Jalapeño: it beats Nvidia Blackwell across almost all scenarios without being tuned for any single point on the curve. Concretely, across the tested range it delivers 1.5x to 1.9x higher peak throughput per watt and 1.7x to 3.6x lower end-to-end latency, winning at both the low-latency and high-throughput ends.
read the caveatsSemiAnalysis walked its own claim back in the same breath, calling the comparison “somewhat incomplete and unfair” because Jalapeño uses newer HBM4 memory, making Nvidia’s Rubin — also on HBM4 — the like-for-like reference. The sharper gap is timing: Rubin is already shipping to customers while Jalapeño remains an engineering sample, not commercially available. The chip itself leaned heavily on AI during its design and was unpacked publicly at Hot Chips 2026.
the cost is in softwareWhether in-house silicon can beat Nvidia is the year’s biggest structural question, and this is the first substantial third-party teardown. Researcher nrehiew_’s first reaction named the practical cost: “time to learn yet another DSL.” Nvidia’s hardest moat was never single-chip performance but a decade-plus of software built on CUDA, which makes developer migration cost the mountain Jalapeño actually has to climb. For Nvidia investors, the thing to track is not this benchmark set but Rubin’s shipping pace racing customers’ own silicon.
▪ SIGNALThe chip that is 1.9x better per watt is an engineering sample; the one already in customer racks is Rubin. The real stake here is time.