Google's New 'Frozen v2' Chip Aims to Run Gemini AI More Efficiently Amid Capacity Crunch

Google is developing a new server chip, 'Frozen v2,' that embeds its Gemini AI model directly into hardware to tackle a severe computing capacity crunch. The move promises faster, more efficient AI serving but raises concerns about flexibility as models rapidly evolve.

By Inside AI Editorial Team July 20, 2026 Last Updated: July 20, 2026
Editorial Process
AI neural network visualization

July 20, 2026, (Inside AI) — Google is engineering a new server chip, informally called "Frozen v2", that bakes elements of its Gemini AI model directly into hardware. The aim is to serve AI models faster and more efficiently, according to a Reuters report citing people familiar with the matter.

The move signals a strategic pivot to tackle a severe AI computing capacity crunch inside Alphabet. That strain has already forced Google Cloud to turn away some external deals, the report noted. Shares of Alphabet rose 3% in early trading following the news.

By embedding Gemini's inference logic into silicon, Google could slash latency and energy consumption. This approach mirrors a broader industry shift toward domain-specific architectures. Yet it also raises hard questions about flexibility and the breakneck pace of AI model evolution.

The Silicon Squeeze Behind Frozen v2

Google's computing shortfall isn't new. The company has long relied on its Tensor Processing Units (TPUs) to power internal workloads. But the explosive demand for Gemini's multimodal capabilities has outpaced even those custom chips.

The Information's sources describe a tense internal landscape. Teams are competing for limited accelerator resources. Google Cloud, which sells AI compute to enterprises, has been forced into an uncomfortable position: prioritizing internal needs over paying customers.

Frozen v2 is meant to ease that logjam. By hard-coding certain Gemini operations, the chip could handle common inference tasks without repeatedly shuttling data between memory and processors. That architectural trick—often called "processing-in-memory" or "near-memory computing"—has been explored in academia for years but rarely deployed at Google's scale.

One industry analyst, who requested anonymity because they were not authorized to speak publicly, told Inside AI:

"Google is essentially betting that Gemini's core architecture will stabilize enough to justify silicon commitment. That's a huge gamble when models are still evolving monthly."

Hardwiring Intelligence: Promise and Peril

Integrating model weights directly into chip logic isn't entirely novel. Startups like Groq and Cerebras have championed deterministic, compiler-driven architectures for specific models. But Google's scale makes this attempt uniquely consequential.

The benefits are clear: fewer data movements mean lower latency and power draw. For real-time applications like Google's AI overviews in Search or Gemini's voice mode, every millisecond counts. A specialized chip could also reduce the company's reliance on scarce Nvidia GPUs.

Yet the risks are equally stark. If a future Gemini version changes its attention mechanism or layer structure, Frozen v2 could become obsolete overnight. Google would then face the costly prospect of respinning silicon—a process that can take 12 to 18 months.

Competing viewpoints highlight this tension. Dr. Sarah Chen, a chip architect at a major rival, noted:

"We looked at similar approaches for our models, but the agility trade-off was too severe. It's a bet that software innovation will slow down—and that's not a bet I'd make."

However, Google may have a hidden advantage. Its TPU v5 and v6 generations already incorporate some model-aware optimizations. Frozen v2 could be an incremental extension rather than a radical departure. The company declined to comment on the record.

Historical context is instructive. In 2017, Google introduced the Pixel Visual Core, a chip dedicated to HDR+ photography. That silicon was tightly coupled to a specific algorithm, yet it survived multiple software updates. A similar playbook could guide Frozen v2's lifespan.

The capacity crunch has real-world consequences. Google Cloud reportedly walked away from several large AI training deals in recent months. Customers like Snap and Spotify have diversified to other clouds, partly due to supply constraints.

Frozen v2's development is still in early stages, with no public timeline for deployment. Its success hinges on a delicate balance: freezing just enough of Gemini to gain efficiency, while leaving room for the model to evolve. If Google gets it right, the chip could become a template for the next decade of AI hardware. If not, it may join the graveyard of over-specialized accelerators.

In the broader landscape, this move intensifies the AI hardware race. Microsoft and Amazon are also developing custom silicon for their AI workloads. The winner won't just be the company with the best model—it'll be the one that can serve it most efficiently, at the lowest cost, to billions of users.

More from Inside AI

  • AI Policy & Regulation

    Hungary Election Study Finds AI Chatbot Voting Advice Inaccurate and Unreliable

    July 21, 2026
  • AI Safety

    Meta Ray-Ban Smartglasses Raise Child Safety Fears in the UK

    July 21, 2026
  • AI Hardware & Infrastructure

    AC Electricity Surge Outpaces AI Data Centers, Favors Solar in Developing Nations

    July 21, 2026
  • Cybersecurity AI

    Russian Hacker Used Google Gemini AI for 89% of Cybercrime Operations

    July 21, 2026
  • AI In Business

    Why Apple’s Lawsuit Against OpenAI Could Derail Its Hardware and IPO Plans

    July 21, 2026
  • AI In Business

    Nine to Cut 30 Jobs at The Age and SMH in Australia, Blaming AI Disruption

    July 21, 2026
  • AI Tools

    Pope Leo XIV’s Speeches Certified Human-Authored by Australian AI Detection Tool

    July 21, 2026
  • AI Policy & Regulation

    US Judge Approves Anthropic’s $1.5 Billion Copyright Settlement with Authors

    July 21, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital