Google's Frozen v2 AI Chip Aims to Embed Gemini in Hardware by 2028

Google is developing 'Frozen v2', a new AI server chip that embeds parts of its Gemini model directly into hardware. The chip promises up to 10x efficiency gains and aims to ease cloud capacity constraints by 2028.

By Inside AI Editorial Team July 21, 2026 Last Updated: July 21, 2026
Editorial Process
AI neural network visualization

July 21, 2026, (Inside AI) — Google is developing a new AI server chip, codenamed 'Frozen v2', designed to embed parts of its Gemini model directly into hardware. The aim is to boost efficiency for AI workloads, particularly inference, according to a report from The Information.

The chip is expected to deliver six to ten times more efficiency than Google's latest custom AI accelerators, measured by AI tokens served per unit of power. Deployment is targeted for 2028, though engineers are still finalizing the design and the extent of hardware integration.

The move comes as Google Cloud faces capacity constraints that have forced it to turn away some external customers. By hardwiring Gemini-specific operations, Frozen v2 could alleviate these bottlenecks and reduce reliance on scarce general-purpose compute.

A Google Cloud spokesperson told The Information: "Our teams are constantly researching and experimenting with new innovations... By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized."

A New Chip Family, Not a TPU Replacement

Frozen v2 is part of a separate chip family, distinct from Google's Tensor Processing Units (TPUs). While TPUs handle both training and inference across diverse models, Frozen chips are tailored for specific inference tasks, likely those central to Gemini's architecture.

This specialization mirrors broader industry trends. Amazon's Inferentia and Microsoft's rumored Athena chip also target inference efficiency. Google's approach, however, ties the hardware directly to its proprietary model, potentially locking in performance gains unavailable to competitors.

The efficiency leap—up to 10x over current TPUs—suggests a radical departure from general-purpose design. It may involve embedding model weights or computational graphs into silicon, a technique known as 'hardwiring' that slashes latency and power draw.

Analysts note that such integration risks inflexibility. If Gemini's architecture evolves, the chip could become obsolete faster than programmable alternatives. Google must balance the gains against the pace of model iteration.

Industry Context and Competitive Pressures

The Frozen project surfaces amid an AI infrastructure spending boom. Global semiconductor investments have soared, driven by demand for large language models. Yet, as Morgan Stanley cautioned in a recent note, expectations for AI chip demand remain "exceptionally high," and firms may struggle to justify ballooning capital expenditures.

Google's custom silicon strategy has already paid dividends. Its TPUs, now in their fifth generation, power internal services like Search and external cloud offerings. But the insatiable appetite for AI compute—exacerbated by models like Gemini—has outpaced supply, leading to the reported customer rejections.

By offloading inference to Frozen v2, Google could free up TPU capacity for training and other tasks. This dual-chip approach echoes Apple's use of Neural Engines alongside GPUs, optimizing for specific workloads without sacrificing versatility.

Competitors are watching closely. Nvidia's dominance in AI chips remains unchallenged, but hyperscalers are increasingly designing their own silicon to reduce costs and dependency. Google's move could pressure Nvidia to offer more tailored solutions or risk losing cloud inference business.

The timeline to 2028 gives Google room to refine the design, but it also leaves a window for rivals. Any delays could erode the advantage, especially if model architectures shift toward sparsity or mixture-of-experts, which may not align with hardwired logic.

Frozen v2's success hinges on execution. Google must navigate the complexities of co-designing hardware and software while maintaining flexibility. If done right, it could set a precedent for model-specific chips, reshaping how AI infrastructure is built.

More from Inside AI

  • AI Policy & Regulation

    US AI Safety Agency Director Resigns After Three Months

    July 20, 2026
  • AI In Business

    Cory Doctorow’s New Book Exposes Who Really Benefits from AI

    July 20, 2026
  • AI Tools

    AWS Launches One-Click Lambda Prompt and OpenAI GPT-5.6 on Bedrock

    July 20, 2026
  • AI Hardware & Infrastructure

    The Hidden Storage Tax Crippling Enterprise AI Conversations

    July 20, 2026
  • AI Hardware & Infrastructure

    Valeo to Make Rare-Earth-Free Drone Motors in France for Harmattan AI

    July 20, 2026
  • AI Hardware & Infrastructure

    Google’s New ‘Frozen v2’ Chip Aims to Run Gemini AI More Efficiently Amid Capacity Crunch

    July 20, 2026
  • Agentic AI

    Google AlphaEvolve AI Agent Automates Code Optimization for Enterprises

    July 20, 2026
  • Artificial Intelligence (AI)

    Bristol Myers Squibb Buys Nvidia’s Latest AI Supercomputer for Drug Research

    July 20, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital