OpenAI's Jalapeño Chip Delivers 1.9x AI Work Per Watt in First Results

OpenAI's first custom inference chip, Jalapeño, delivers industry-leading throughput and latency gains across multiple AI models, signaling a new era in AI hardware.

Last Updated: September 13, 2026 Editorial Process
Editorial Process
See more of Inside AI's trusted news by adding us as a preferred source on Google.
AI neural network visualization
Published on: August 25, 2026

August 25, 2026, (Inside AI) — OpenAI has disclosed the first measured results for its custom inference chip, Jalapeño, showing a 1.5 to 1.9 times increase in AI work per watt at peak throughput compared to leading commercial systems. The chip also cut end-to-end latency by 1.7 to 3.6 times across three public models: GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T.

The announcement marks a shift in the AI hardware race. For years, inference accelerators forced a tradeoff between throughput and latency. Jalapeño appears to break that pattern by delivering both in a single architecture, a feat that could reshape how interactive AI agents are deployed at scale.

OpenAI said the chip is rated at 700 watts, but measured sustained power stayed at or below 550 watts on tested workloads. That efficiency gain matters because AI inference costs are now a dominant factor in the economics of large language models.

Why Agents Expose the Limits of Current Inference Hardware

Agentic workloads are uniquely punishing for inference systems. An agent may need dozens of sequential model calls to complete one task. If each call adds 100 milliseconds of latency, the total delay compounds into seconds of perceived unresponsiveness.

Read: OpenAI and Samsung Deepen Partnership on Next-Generation AI Chips

Jalapeño was designed specifically for this problem. The chip minimizes data movement and communication delays by keeping model state, including the KV cache, local to the compute units that need it. The network is integral to the architecture, allowing an entire workload to remain within one connected system.

Language-model inference has two distinct phases with different bottlenecks. Prefill, when the system processes a prompt, is compute-intensive. Decode, when the system generates tokens one by one, is constrained by memory bandwidth. A system that excels at one phase can lose that advantage while waiting for data to move between cores and chips.

Jalapeño addresses this by explicitly placing model state and activating the right combination of compute, memory, and networking for each phase. The result is a balanced accelerator that can adapt as the balance between prefill and decode changes, a defining feature of agentic workloads.

AI Helped Design the Chip That Runs AI

OpenAI said AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months. AI explored implementations, shortened design and verification loops, and optimized arithmetic circuits to fit more compute performance into the chip on schedule.

The company also used its Codex model with GPT‑Astra to bring three open-weight models that were not part of the original production plan to high performance within two months. For selected GPT‑OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than existing human-expert-written implementations.

Those figures apply only to selected blocks, not the full model, but they point to a powerful new development loop. Engineers can describe work through local tensors, explicit communication, and predictable synchronization. AI can then optimize how that work is mapped, placed, scheduled, and coordinated across the system.

Read: Microsoft Plans Maia 300 AI Chip Unveiling in September, Report Says

OpenAI tested Jalapeño on InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request. The company compared Jalapeño with leading commercially available AI systems across the tested operating range, from high-throughput serving to highly interactive, low-latency use.

On Kimi K2.5 1T, the largest public model tested, Jalapeño delivered approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system. In internal testing, the advantage widened further on frontier OpenAI models, suggesting the architecture becomes more valuable as workloads grow larger.

OpenAI plans to begin deploying Jalapeño within its compute infrastructure by the end of the year. The company said Gen 2 is deep in development and Gen 3 is taking shape. Each generation will build on what the team learns and further advance both efficiency and speed.

Meeting growing demand for AI will require more compute from every available source. OpenAI said it will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads. The mission, the company said, is to ensure that artificial general intelligence benefits all of humanity.

More from Inside AI

  • Features, Interviews, Press Releases

    Beyond Transcripts: Modulate Secures $25M to Scale Frontier Audio-Native AI Architecture Against Monolithic LLMs

    September 28, 2026
  • AI Safety

    Anthropic AI Model Sent Fake Homicide Tip to Philadelphia Police

    October 10, 2026
  • AI In Business

    OpenAI’s September Revenue Falls Short of Internal Targets

    October 9, 2026
  • Agentic AI

    Meta’s Muse vs Amazon: Who Owns the Customer in the AI Agent Era?

    October 9, 2026
  • AI Safety

    OpenAI defends firing of 3 AI safety researchers, cites ‘significant breach of trust’

    October 9, 2026
  • AI In Business

    AI Companies to Drive Most Third-Quarter US Earnings Gains, LSEG Data Shows

    October 9, 2026
  • AI In Business

    AI-Driven US Bull Market Nears Four-Year Anniversary as Concentration Risks Grow

    October 9, 2026
  • Agentic AI

    Google Cloud Launches Gemini Agent for Workplace AI Tasks

    October 9, 2026
  • AI Hardware & Infrastructure

    Infino AI Raises $7.5M Seed to Unify Data Retrieval for AI Agents

    October 9, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Join Our Newsletter Community

Subscribe

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Policy & Regulation
  • AI Hardware & Infrastructure
  • AI Tools
  • AI In Business
  • Robotics
  • Cybersecurity AI
  • AI Safety
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases
  • Features
  • Sponsored Content
  • Advertise with us
  • Newsletter

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital