Anthropic has committed to buying up to one million of Google's Ironwood TPUs, a show of confidence as Google doubles down on custom silicon to cut reliance on Nvidia GPUs. Google unveiled two new custom Tensor Processing Units this week — rolling them into Google Cloud and pairing them with a multi‑partner supply chain that includes Broadcom, MediaTek and Marvell — as part of a broader push to lower the cost of AI inference.

Two chips, one strategy

Google's latest hardware push centers on a pair of custom TPUs intended to handle the heavy lifting of modern AI workloads while avoiding some of the bottlenecks and price pressure tied to Nvidia GPUs. Alphabet introduced the new TPUs this week as part of a broader shift toward so‑called silicon independence among hyperscalers, with an explicit goal of lowering the cost of inference for enterprise customers.

Ironwood — Google's seventh‑generation TPU — is already live on Google Cloud and designed specifically for inference workloads. The chip delivers roughly ten times the peak performance of Google's prior v5p accelerators, offers 192 gigabytes of HBM3E memory per chip and a memory bandwidth of about 7.2 terabytes per second, according to product details shared by the company. Google says Ironwood can scale to 9,216 liquid‑cooled chips in a single superpod that can produce 42.5 FP8 exaflops.

But Ironwood is only part of the picture. Google is laying out a split roadmap for the TPU v8 generation: one variant targeting training and high‑performance tasks, and another optimized for inference and lower cost. The training chip is being designed by Broadcom under the codename Sunfish, while MediaTek is handling the inference variant codenamed Zebrafish.

A diversified supply chain

Google has intentionally broadened the list of partners behind its silicon effort. Broadcom, MediaTek, Marvell and Intel now all play defined roles in the roadmap. Broadcom signed a long‑term supply agreement earlier this year to provide TPUs and networking components through 2031 and is tasked with the high‑end training silicon. MediaTek is building cost‑optimized inference parts and already contributed I/O modules and peripheral designs for Ironwood.

MediaTek's designs reportedly run 20 to 30 percent cheaper than alternative approaches in the inference segment, a price gap Google sees as material when serving large‑scale cloud customers. Marvell is in talks with Google to add a memory processing unit and an additional inference TPU to the family, a development that could shift some memory‑bound work off the main accelerator and reduce latency for memory‑heavy models.

That multi‑vendor approach gives Google negotiating leverage with foundries and suppliers and reduces single‑vendor exposure. It also mirrors strategies at other hyperscalers that have moved to combine internal accelerators with third‑party GPUs, depending on workload and customer choice.

Technical trade‑offs and the memory angle

The memory processing unit Marvell is discussing would handle data movement and memory operations separate from the main TPU. Offloading those tasks could ease one of the key bottlenecks in modern AI stacks: feeding large models with steady, low‑latency access to weights and activation data. Separating memory functions from compute could cut latency and improve overall throughput for some inference workloads.

Google's Ironwood already packs high‑bandwidth memory and is tuned for inference density, but adding a dedicated MPU to the platform would let the company tune system‑level trade‑offs more finely — pushing more memory work to a specialist chip and keeping heavyweight matrix math on the TPU itself.

Memory costs and supply tightness have become a constraint for large‑scale inference deployments. By designing elements of the memory path, Google can better tune system trade‑offs, control costs and ease supply constraints for large‑scale inference deployments.

Related Articles

Anthropic has committed to buying up to one million Ironwood TPUs, and Google expects to produce millions of units this year.

This article was created with AI assistance.