
Jalapeño’s first results show industry-leading speed and efficiency in AI inference
OpenAI developed Jalapeño, its first custom inference chip designed to serve AI requests with greater speed and power efficiency.
Why it matters
This could result in faster AI responses and more responsive agents for users. It may also make powerful AI more affordable and widely available by reducing operational costs.
The details
- Deployment within OpenAI's compute infrastructure is planned by the end of 2026.
- Jalapeño achieved 1.7 to 3.6 times lower end-to-end latency than comparison systems.
- The chip was developed in nine months using AI for design and optimization.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.