July 21, 2026, (Inside AI) — Google is developing a new AI server chip, codenamed 'Frozen v2', designed to embed parts of its Gemini model directly into hardware. The aim is to boost efficiency for AI workloads, particularly inference, according to a report from The Information.
The chip is expected to deliver six to ten times more efficiency than Google's latest custom AI accelerators, measured by AI tokens served per unit of power. Deployment is targeted for 2028, though engineers are still finalizing the design and the extent of hardware integration.
The move comes as Google Cloud faces capacity constraints that have forced it to turn away some external customers. By hardwiring Gemini-specific operations, Frozen v2 could alleviate these bottlenecks and reduce reliance on scarce general-purpose compute.
A Google Cloud spokesperson told The Information: "Our teams are constantly researching and experimenting with new innovations... By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized."
A New Chip Family, Not a TPU Replacement
Frozen v2 is part of a separate chip family, distinct from Google's Tensor Processing Units (TPUs). While TPUs handle both training and inference across diverse models, Frozen chips are tailored for specific inference tasks, likely those central to Gemini's architecture.
This specialization mirrors broader industry trends. Amazon's Inferentia and Microsoft's rumored Athena chip also target inference efficiency. Google's approach, however, ties the hardware directly to its proprietary model, potentially locking in performance gains unavailable to competitors.
The efficiency leap—up to 10x over current TPUs—suggests a radical departure from general-purpose design. It may involve embedding model weights or computational graphs into silicon, a technique known as 'hardwiring' that slashes latency and power draw.
Analysts note that such integration risks inflexibility. If Gemini's architecture evolves, the chip could become obsolete faster than programmable alternatives. Google must balance the gains against the pace of model iteration.
Industry Context and Competitive Pressures
The Frozen project surfaces amid an AI infrastructure spending boom. Global semiconductor investments have soared, driven by demand for large language models. Yet, as Morgan Stanley cautioned in a recent note, expectations for AI chip demand remain "exceptionally high," and firms may struggle to justify ballooning capital expenditures.
Google's custom silicon strategy has already paid dividends. Its TPUs, now in their fifth generation, power internal services like Search and external cloud offerings. But the insatiable appetite for AI compute—exacerbated by models like Gemini—has outpaced supply, leading to the reported customer rejections.
By offloading inference to Frozen v2, Google could free up TPU capacity for training and other tasks. This dual-chip approach echoes Apple's use of Neural Engines alongside GPUs, optimizing for specific workloads without sacrificing versatility.
Competitors are watching closely. Nvidia's dominance in AI chips remains unchallenged, but hyperscalers are increasingly designing their own silicon to reduce costs and dependency. Google's move could pressure Nvidia to offer more tailored solutions or risk losing cloud inference business.
The timeline to 2028 gives Google room to refine the design, but it also leaves a window for rivals. Any delays could erode the advantage, especially if model architectures shift toward sparsity or mixture-of-experts, which may not align with hardwired logic.
Frozen v2's success hinges on execution. Google must navigate the complexities of co-designing hardware and software while maintaining flexibility. If done right, it could set a precedent for model-specific chips, reshaping how AI infrastructure is built.