“We made a chip, and it is fast.”
With that simple, understated post on X, OpenAI CEO Sam Altman officially marked a massive turning point for the AI giant. The company behind ChatGPT is currently building the very hardware their AI models run on.

Developed in partnership with Broadcom and playfully dubbed “Jalapeño,” OpenAI’s first custom artificial intelligence processor has arrived. The chip is already posting benchmark results that are turning heads across the industry.
Here is a breakdown of what Jalapeño is, how it performs, and why it matters for the future of AI.
Built for Inference, Not Training
To understand Jalapeño, you have to understand the two main phases of an AI model’s lifecycle: training and inference.
Training is the heavy, glamorous work of teaching an AI model from scratch, a process that requires tens of thousands of GPUs running for months. Jalapeño is not built for this.
Instead, Jalapeño is purpose-built for inference, the unglamorous but incredibly expensive job of actually answering your prompts, writing code, and generating responses once the model already exists. Every time a user interacts with ChatGPT or Codex, an inference chip is doing the work. By building a processor that does exactly one thing at massive scale, OpenAI aims to shift its unit economics and power consumption fundamentally.
Unprecedented Speed and Efficiency
According to recent benchmarks verified by SemiAnalysis, Jalapeño is a major technological leap. When tested on the InferenceX benchmark against current-generation Nvidia systems (like the GB200 and GB300), OpenAI’s new silicon delivered impressive numbers:
| Metric | Performance vs. Leading Competitors |
| Power Efficiency | 1.5x to 1.9x more AI work per watt |
| Response Latency | 1.7x to 3.6x faster end-to-end response times |
| Highly Interactive Workloads | 2.1x to 4.1x higher performance |
| Sustained Power | Rated for 700W, but operated efficiently below 550W during tests |
OpenAI’s Vice President of Hardware, Richard Ho, noted that standard AI chips usually force a trade-off between throughput (handling lots of requests at once) and latency (speed of a single response). Jalapeño’s custom architecture circumvents this by minimising data movement, keeping the model’s key-value cache locally on the chip.
What This Means for Nvidia (and the Industry)
Nvidia has enjoyed a near-monopoly on the AI boom, supplying the vast majority of the world’s highly coveted GPUs. OpenAI has historically been one of Nvidia’s biggest and most dependent customers.
Jalapeño represents a major shift toward OpenAI becoming a “full-stack” AI company. While OpenAI has confirmed they will continue relying on Nvidia chips for training and as a compute partner, moving inference workloads to in-house silicon will drastically reduce their reliance on a single supplier. They are joining the ranks of Google (TPUs), Amazon (Trainium/Inferentia), and Microsoft (Maia), all of which have invested heavily in proprietary silicon.
Looking Ahead
OpenAI plans to deploy Jalapeño in small volumes by the end of 2026, with a massive scale-up planned for 2027. And they aren’t stopping there: Generation 2 of the chip is already “deep into development,” and Generation 3 is taking shape.
As models get smarter, the physical infrastructure running them must keep pace. With Jalapeño, OpenAI is proving they intend to control their destiny down to the very silicon their AI calls home.



