Google’s ‘Frozen v2’ AI Chip Could Change the Economics of Gemini

Google's 'Frozen v2' AI Chip Could Change the Economics of Gemini

Google is reportedly working on an ambitious new AI processor that could fundamentally change how its Gemini models are deployed. Internally known as “Frozen v2,” the custom server chip is designed to embed parts of Gemini directly into hardware, allowing AI responses to be generated far more efficiently than today’s software-only approach.

If successful, it would represent the next phase of AI infrastructure—where the biggest innovation is no longer just better models, but purpose-built silicon optimized for a single AI family.

Why Google Needs a New AI Chip

The AI race has created an unprecedented demand for computing power.

Training frontier AI models requires enormous GPU clusters, while serving billions of daily AI queries consumes even more electricity and hardware capacity. Reports suggest Google Cloud has already faced capacity constraints severe enough to decline some enterprise deals.

Rather than relying solely on increasingly expensive AI accelerators, Google is pursuing hardware-software co-design—building chips specifically around Gemini’s architecture.

This is the same philosophy that made Apple’s M-series processors highly efficient: optimize hardware for the software it runs most often.

What Is “Frozen v2”?

According to reports, Frozen v2 will not replace Google’s Tensor Processing Units (TPUs). Instead, it will complement them by hardwiring portions of Gemini’s neural network into silicon.

Unlike programmable AI chips that execute every model instruction dynamically, Frozen v2 would permanently encode selected parts of Gemini into hardware. This reduces computation overhead, lowers latency, and dramatically improves energy efficiency.

Google reportedly aims to deploy the chip around 2028, although engineers are still finalizing its design.

Google’s Long History of AI Chips

Google has been building custom AI chips long before the generative AI boom.

  • 2016: Google introduced the first Tensor Processing Unit (TPU) to accelerate machine learning workloads.
  • 2021–2025: Successive TPU generations powered Search, YouTube, Google Cloud, and Gemini training.
  • 2025: Google unveiled Ironwood TPU, designed specifically for large-scale AI inference, delivering massive gains in performance and energy efficiency.
  • 2026: Reports now suggest Google is exploring an entirely new class of AI processors through the Frozen v2 project.

Rather than replacing TPUs, Frozen v2 would work alongside them, focusing on serving Gemini responses faster and at dramatically lower power consumption.

Up to 10× Better Efficiency

One of the most striking claims is efficiency.

Reports suggest Frozen v2 could deliver 6–10 times more AI tokens per unit of power compared with Google’s latest AI hardware.

For Google, this translates into several strategic advantages:

  • Lower operating costs for Gemini
  • Faster AI responses
  • Reduced electricity consumption
  • Greater AI capacity without building proportionally larger data centers

As AI inference becomes the dominant cost of running large language models, efficiency may become just as important as intelligence.

A Response to Fierce AI Competition

The timing is notable.

Google is under intense pressure from OpenAI, Anthropic and rapidly advancing Chinese AI companies such as Moonshot AI and DeepSeek.

Last week, reports indicated that Gemini 3.5 Pro had been delayed after falling short of Google’s internal performance goals, particularly in coding.

Frozen v2 shows that Google is pursuing a two-pronged strategy:

  • Improve Gemini’s capabilities.
  • Reduce the cost of delivering those capabilities at global scale.

Winning the AI race increasingly depends on both.

Why Custom AI Silicon Is the New Battleground

The industry’s biggest AI players are no longer competing only with models.

Microsoft relies heavily on NVIDIA GPUs.

Amazon has Trainium and Inferentia.

Meta is developing custom AI accelerators.

Google already operates one of the world’s largest TPU fleets and now appears ready to take integration even further.

The future competitive advantage may belong to companies controlling the entire AI stack—from silicon and networking to models and cloud infrastructure.

The Bigger Picture

Frozen v2 reflects a broader shift in artificial intelligence. The next breakthroughs may come less from adding trillions of new parameters and more from making existing models dramatically cheaper and faster to run.

If Google’s reported efficiency gains materialize, the company could reduce the cost of serving billions of Gemini requests while easing pressure on its AI infrastructure. In an era where compute has become the world’s most valuable AI resource, custom silicon may prove as important as the models themselves.