OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new system. Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors.

By AI NewsroomPublished about 5 hours agoUpdated about 1 hour ago4 views
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

Why It Matters

This story touches on openai, jalapen, more — topics readers are actively tracking. Review and add editorial context before publishing.

Key Facts

  • Fact 1: Ho estimated that Jalapeño would deploy at the end of 2026 “in very small volumes,” with more significant deployment coming in 2027.
  • Fact 2: Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies.
  • Fact 3: He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489.

At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new system. Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors.

“The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardware, in a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.” Notably, that comparison is against an Nvidia Blackwell system — but by the time Jalapeño reaches full deployment, the competition may have advanced significantly.

(Original synthesis pending human/AI review — generated by the stub provider by selecting real sentences from the source material, not by writing new analysis or commentary.)

Original source: TechCrunch

Share

Related Stories