Dynamic Beating AI News Flash: OpenAI has released the latest test results of its self-developed AI inference chip, Jalapeño. The chip, named after the Mexican pepper, outperformed the NVIDIA GB200 and GB300 superchips in the InferenceX benchmark. For models such as GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, the AI workload per watt achieved 1.5 to 1.9 times that of competitors, with end-to-end latency reduced by 1.7 to 3.6 times.
OpenAI stated that Jalapeño uses an ASIC architecture developed in collaboration with Broadcom, specifically designed for AI inference to increase throughput and reduce latency simultaneously. OpenAI plans a small-scale deployment by the end of this year, with expanded production in 2027, and ongoing development of second and third-generation chips. However, the company emphasized that these chips will not entirely replace those from partners like NVIDIA.

