NVIDIA Groq 3 LPX Enters Full Production for Agentic AI Inference

NVIDIA's Groq 3 LPX accelerator is now in full production, delivering record token generation speeds for agentic AI, with Nebius as the first cloud adopter.

Last Updated: August 24, 2026 Editorial Process
Editorial Process
See more of Inside AI's trusted news by adding us as a preferred source on Google.
AI neural network visualization
Published on: August 24, 2026

August 24, 2026, (Inside AI) — NVIDIA has moved its Groq 3 LPX inference accelerator into full production, targeting the explosive token demands of agentic AI. The chip is an extension of the Vera Rubin platform and is built to slash latency for interactive workloads.

Nebius is the first AI cloud to adopt the accelerator, integrating it into its Token Factory production inference service. The announcement came at the Hot Chips conference, where NVIDIA framed the move as a direct response to the economics of agentic computing.

Agentic systems can generate massive volumes of tokens across hundreds or thousands of inference steps. Faster token generation is critical for agents to reason, act, and complete complex tasks in real time. Groq 3 LPX is purpose-built to extend Vera Rubin's interactivity, the rate at which tokens are generated for an individual user.

In Artificial Analysis benchmarking, the accelerator delivered a record 3,400 output tokens per second running Gemma 4 31B, an open source agentic model. The test used a 100,000-token context, a requirement for complex agent workflows. NVIDIA said this is the fastest performance ever recorded for the model.

The company claims 4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform. For coding agents, that translates to tasks completed in minutes instead of hours, with more time to inspect files, write and test code, call tools, verify results, and iterate.

Why token generation speed now defines the inference market

The shift from static chatbots to autonomous agents has changed what matters in inference hardware. A single agentic task can require hundreds of sequential model calls, each dependent on the previous output. Total task time is dominated by per-token latency, not raw throughput.

This is why NVIDIA is positioning Groq 3 LPX as an interactive accelerator rather than a general-purpose inference chip. The distinction is technical and commercial. Traditional inference accelerators optimize for tokens per second per dollar across a batch of users. Interactive accelerators optimize for tokens per second per individual user, the metric that determines whether an agent feels responsive.

NVIDIA's Grace Blackwell and NVL72 systems already delivered a major leap in large language model inference. Vera Rubin extends that vision with workload-optimized AI factory configurations. Groq 3 LPX is the interactive tier of that platform.

"Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency," said Jensen Huang, founder and CEO of NVIDIA. "Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation. This transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness, just as demand for AI computation is accelerating worldwide."

The competitive landscape is fragmented. Startups like Groq, the inference cloud company that shares a name with the accelerator but is a separate entity, have built businesses around low-latency token generation. Hyperscalers are deploying custom silicon. NVIDIA's move signals that it sees interactive inference as a distinct market segment worth defending with dedicated hardware.

Nebius bets on instant agent loops without stack migration

Nebius plans to offer Groq 3 LPX through its Token Factory platform, which already serves production inference workloads. The pitch to developers is simple: faster token generation without changing APIs or migrating to a new stack.

"Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what NVIDIA Groq 3 LPX is built to accelerate," said Danila Shtan, chief technology officer of Nebius. "As the first AI cloud bringing it to production via Nebius Token Factory, we're making sure every step of an agent's loop feels instant -- through the same API developers are already using, with no migration to a new stack."

The adoption pattern matters. AI clouds have become the engines of the AI economy, giving enterprises and developers access to advanced infrastructure without owning hardware. For providers serving latency-sensitive, high-volume inference workloads, Groq 3 LPX offers a path to deploy differentiated compute in proven rack-scale systems.

The purpose-built AI inference cloud Groq also plans to be among the platform's earliest adopters, following Nebius. That creates an unusual situation where a company named Groq will operate NVIDIA hardware called Groq 3 LPX, a naming collision that has already caused confusion in industry discussions.

NVIDIA's Vera Rubin platform uses extreme codesign across seven chips and five purpose-built racks. The NVL72 and Groq 3 LPX configurations tackle different workload requirements, from frontier model training to open model serving. The racks feature BlueField-4 DPUs and work with Vera CPU racks, Vera BlueField-4 STX storage, and Spectrum-6 SPX Ethernet to optimize multi-agent systems for throughput per watt and low-latency inference.

The broader industry context is a race to make agents commercially viable. If an agent takes minutes to complete a task that a human does in seconds, adoption stalls. Hardware that reduces per-token latency directly addresses that bottleneck. NVIDIA's full production status for Groq 3 LPX suggests it believes the agentic AI market is now real enough to justify dedicated silicon.

More from Inside AI

  • AI Hardware & Infrastructure

    NVIDIA Groq 3 LPX Enters Full Production for Agentic AI Inference

    August 24, 2026
  • AI In Business

    IIT Kharagpur Outreach Network Launches Texas AI Learning Centre in US

    August 24, 2026
  • AI In Business

    OpenAI GPT-5.6 Terra and Luna Now Available on Amazon Bedrock in AWS GovCloud

    August 24, 2026
  • AI Hardware & Infrastructure

    Trump-linked RUM Group signs $13.7 billion AI chip deal with US cloud client

    August 24, 2026
  • Generative AI

    Alibaba launches Wan3.0 AI video model after $10 billion share sale

    August 24, 2026
  • AI In Business

    S&P, Nasdaq Futures Slip as Markets Await Iran Sanctions, Nvidia Results

    August 24, 2026
  • AI In Business

    Sonova unveils new AI-powered hearing-aid platform to support growth targets

    August 24, 2026
  • Robotics

    ENGINEAI Says General-Purpose Humanoid Robot Costs Fell Below RMB100,000

    August 24, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Policy & Regulation
  • AI Hardware & Infrastructure
  • AI Tools
  • AI In Business
  • Robotics
  • Cybersecurity AI
  • AI Safety
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases
  • Features
  • Sponsored Content

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital