Anthropic plans to provision up to one million of Google’s new Ironwood TPUs — a strong signal that customers are betting on inference-optimized silicon to cut AI latency and costs. Google Cloud introduced Ironwood, its seventh-generation TPU, at Cloud Next in April 2025; a single liquid-cooled pod can link as many as 9,216 chips through an inter‑chip network that draws nearly 10 megawatts.

Ironwood: a shift toward inference

Google unveiled Ironwood as its most powerful TPU yet and pitched it as the first TPU explicitly designed for inference, not just training. The company introduced the chip at Google Cloud Next in April 2025 and has been preparing a wider public roll-out since, with commercial availability scheduled in the weeks after that announcement.

Ironwood can be liquid-cooled and connected at scale: a single pod can link up to 9,216 chips through an Inter-Chip Interconnect network that Google says spans nearly 10 megawatts. The setup is part of what Google calls its Cloud AI Hypercomputer architecture, which pairs hardware and software to handle both demanding training workloads and very fast serving of models.

On the software side, Google points customers to its Pathways stack, which is meant to let developers tap the combined power of many Ironwood TPUs reliably. The company says Ironwood is more energy efficient and more performant than previous TPU generations; Google has described it as offering multiple times the speed of its predecessor for certain workloads.

"We're designing for a new phase of AI where models don't just answer queries but generate insights proactively," the Google blog announcing Ironwood said, framing the product as a step into what the company calls the "age of inference."

How the chips change the trade-offs

For years, most large models and enterprise AI work ran on graphics processing units, or GPUs, with Nvidia the dominant supplier. Custom silicon like TPUs offers a different set of trade-offs: firms can tailor chips for particular workloads and, in some cases, win on price, efficiency and latency.

  • Specialization: Chips can be optimized for training or for inference, making it sensible to dedicate hardware to each phase.
  • Latency and throughput: Inference-focused designs prioritize low latency and predictable throughput to serve many simultaneous real-time requests.
  • Training needs: By contrast, training emphasizes raw compute and memory bandwidth.

"It now becomes sensible to specialize chips more for training or more for inference workloads," said Jeff Dean, Google Chief Scientist, in an interview. He argued that dedicating hardware to inference can cut the time it takes to get results when users are querying chatbots and AI agents.

Demis Hassabis, chief executive officer of Google DeepMind, noted the appeal of running on both TPUs and GPUs. "A lot of people would like to run on both," he said, describing why labs and developers value flexibility in hardware choices.

Customer demand and business stakes

Cloud customers have taken notice. Google said Anthropic plans to use up to 1 million new TPUs to power its Claude model, a signal of how leading AI labs are buying custom capacity. Large AI developers and some rivals are stocking up on Google's chips even as they continue to run workloads on GPUs elsewhere.

Sundar Pichai, chief executive officer of Alphabet Inc., told investors the company is seeing strong demand for both TPU- and GPU-based infrastructure and is investing to meet it. In its third-quarter 2025 earnings report, Google reported cloud revenue of US$15.15 billion, a 34% year‑over‑year increase, and raised the high end of its capital-spending forecast to US$93 billion to expand capacity.

Those moves reflect a broader commercial bet on AI infrastructure.

Related Articles

Sundar Pichai said Google is "seeing substantial demand for our AI infrastructure products, including TPU-based and GPU-based solutions," and the company is investing to meet that demand.

This article was created with AI assistance.