July 20, 2026, (Inside AI) — Google is engineering a new server chip, informally called "Frozen v2", that bakes elements of its Gemini AI model directly into hardware. The aim is to serve AI models faster and more efficiently, according to a Reuters report citing people familiar with the matter.
The move signals a strategic pivot to tackle a severe AI computing capacity crunch inside Alphabet. That strain has already forced Google Cloud to turn away some external deals, the report noted. Shares of Alphabet rose 3% in early trading following the news.
By embedding Gemini's inference logic into silicon, Google could slash latency and energy consumption. This approach mirrors a broader industry shift toward domain-specific architectures. Yet it also raises hard questions about flexibility and the breakneck pace of AI model evolution.
The Silicon Squeeze Behind Frozen v2
Google's computing shortfall isn't new. The company has long relied on its Tensor Processing Units (TPUs) to power internal workloads. But the explosive demand for Gemini's multimodal capabilities has outpaced even those custom chips.
The Information's sources describe a tense internal landscape. Teams are competing for limited accelerator resources. Google Cloud, which sells AI compute to enterprises, has been forced into an uncomfortable position: prioritizing internal needs over paying customers.
Frozen v2 is meant to ease that logjam. By hard-coding certain Gemini operations, the chip could handle common inference tasks without repeatedly shuttling data between memory and processors. That architectural trick—often called "processing-in-memory" or "near-memory computing"—has been explored in academia for years but rarely deployed at Google's scale.
One industry analyst, who requested anonymity because they were not authorized to speak publicly, told Inside AI:
"Google is essentially betting that Gemini's core architecture will stabilize enough to justify silicon commitment. That's a huge gamble when models are still evolving monthly."
Hardwiring Intelligence: Promise and Peril
Integrating model weights directly into chip logic isn't entirely novel. Startups like Groq and Cerebras have championed deterministic, compiler-driven architectures for specific models. But Google's scale makes this attempt uniquely consequential.
The benefits are clear: fewer data movements mean lower latency and power draw. For real-time applications like Google's AI overviews in Search or Gemini's voice mode, every millisecond counts. A specialized chip could also reduce the company's reliance on scarce Nvidia GPUs.
Yet the risks are equally stark. If a future Gemini version changes its attention mechanism or layer structure, Frozen v2 could become obsolete overnight. Google would then face the costly prospect of respinning silicon—a process that can take 12 to 18 months.
Competing viewpoints highlight this tension. Dr. Sarah Chen, a chip architect at a major rival, noted:
"We looked at similar approaches for our models, but the agility trade-off was too severe. It's a bet that software innovation will slow down—and that's not a bet I'd make."
However, Google may have a hidden advantage. Its TPU v5 and v6 generations already incorporate some model-aware optimizations. Frozen v2 could be an incremental extension rather than a radical departure. The company declined to comment on the record.
Historical context is instructive. In 2017, Google introduced the Pixel Visual Core, a chip dedicated to HDR+ photography. That silicon was tightly coupled to a specific algorithm, yet it survived multiple software updates. A similar playbook could guide Frozen v2's lifespan.
The capacity crunch has real-world consequences. Google Cloud reportedly walked away from several large AI training deals in recent months. Customers like Snap and Spotify have diversified to other clouds, partly due to supply constraints.
Frozen v2's development is still in early stages, with no public timeline for deployment. Its success hinges on a delicate balance: freezing just enough of Gemini to gain efficiency, while leaving room for the model to evolve. If Google gets it right, the chip could become a template for the next decade of AI hardware. If not, it may join the graveyard of over-specialized accelerators.
In the broader landscape, this move intensifies the AI hardware race. Microsoft and Amazon are also developing custom silicon for their AI workloads. The winner won't just be the company with the best model—it'll be the one that can serve it most efficiently, at the lowest cost, to billions of users.