Google is reportedly building a server chip designed to do one thing: run its Gemini models. Codenamed Frozen v2, it takes the model's internal structure and bakes it directly into silicon, and engineers expect it to deliver 6 to 10 times more performance per watt than the company's current TPUs. Deployment in Google's data centers is said to begin in 2028.
From General-Purpose Silicon to a Chip Built for One Model
AI silicon has so far been built to keep some flexibility. Google's TPUs, offered through its cloud for years, are tuned separately for training and inference, but they are still designed to run a wide range of models. Flexibility is a strength, and it is also a cost: for any single model, part of the chip is circuitry that model never uses.
Frozen v2 drops that assumption. The Gemini architecture, meaning the blueprint of which calculations run in what order, gets fixed into the hardware itself. The chip gives up generality in exchange for stripping almost all of the waste out of running one model. The name Frozen describes the design fairly literally.
The original idea is reported to have come from Jeff Dean, chief scientist at Google DeepMind.
Why Google Froze the Architecture Rather Than the Weights
The more interesting detail is what Google reportedly abandoned. An earlier approach would have burned the model's weights, the parameters produced by training, into the chip. That version would only have worked with one specific release of Gemini and would have aged out the moment the model was updated. Silicon takes years to design and manufacture, which makes that a fatal constraint.
So the thing being frozen shifted from the weights, which change often, to the architecture, which changes far less. The underlying bet is that Gemini's skeleton will not be redesigned every generation. That bet is also the main thing worth watching, because the whole design depends on it holding.
Where the Efficiency Actually Comes From
Two mechanisms account for most of the claimed 6 to 10 times gain.
The first is cutting the number of calculations. Merging several operations into a single computation is a well-established way to speed up AI models, but doing it in the circuitry rather than in software each time means nothing gets left on the table.
The second is reducing data movement. When a model does not fit in on-chip memory, the hardware has to shuttle pieces back and forth with external memory, and that traffic is what slows things down. Give Frozen v2 enough memory to hold Gemini entirely on-chip and the round trip disappears. Moving data, rather than performing arithmetic, is often the heavier part of AI power consumption, so this is not a minor lever.
What to Make of a 2028 Timeline
Alphabet shares rose 1.5 percent on the report. Even so, deployment starts in 2028, and production volume is expected to be smaller than the TPU line. The chip reads less like a volume product and more like a test of whether the specialized-silicon approach pays for itself at all.
Google has not confirmed Frozen v2, and the 6 to 10 times figure has not been measured on shipping silicon. Projections of this kind routinely shrink once a chip reaches production, so taking the number at face value would be premature. Still, with the cost of running inference weighing on every lab, designing model and chip as one piece is a reasonable direction. Google is currently late with Gemini 3.5 Pro, but an advantage in the cost of serving would let it compete on ground that has nothing to do with benchmark wins.
Summary
Frozen v2, the chip Google is reported to be developing, aims for 6 to 10 times better performance per watt by fixing Gemini's architecture into silicon. Freezing the structure rather than the weights avoids obsolescence, and the design targets both fewer calculations and less data movement. Data center rollout is said to start in 2028 at a smaller production scale than TPUs. Google has not confirmed any of it and the numbers remain unverified, but the project is a useful marker for how far the industry will take co-designed models and hardware.
