According to TrendForce's Beating report, Google has been exposed to be developing a Frozen v2 AI Inference Chip. It will directly hardcode part of the Gemini architecture into the chip, reducing computation and data movement when running models.
Internally, it is expected that the number of Tokens processed per watt will reach 6 to 10 times that of the latest TPU. Google plans to deploy it as early as 2028.
Google hopes to use it to alleviate the shortage of computing power. The gap has already forced Google Cloud to reject some external customer orders.
Frozen v2 will coexist with TPUs. TPUs can run different models, while Frozen v2 trades flexibility for efficiency.
This is also an architectural bet. If Gemini undergoes a major redesign in the future, these chips may no longer be usable.
