Google Splits Its AI Chip Into Two: How the TPU 8t and 8i Are Built for the Agentic Era

News Summary
At Google Cloud Next 2026 (April 22–24, Las Vegas, PDT), Google unveiled the eighth generation of its Tensor Processing Units (TPUs), introducing two purpose-built chips — the TPU 8t for AI model training and the TPU 8i for inference — marking the first time Google has separated these workloads into distinct architectures. The TPU 8t delivers nearly 3× the compute performance of its predecessor, while the TPU 8i is engineered for ultra-low latency in agentic and multi-step AI workflows. Both chips are designed as part of Google's broader AI Hypercomputer platform and are expected to be generally available later in 2026.
A New Architecture for the Agentic Era
Google's decision to split its eighth-generation TPU into two specialized processors reflects a fundamental shift in how AI infrastructure is being designed. For over a decade, a single TPU handled both the intensive computation of training large models and the ongoing work of serving those models to end users. As AI systems evolve from simple question-answering tools toward autonomous, multi-step agents capable of executing complex workflows, the performance requirements for training and inference have diverged sharply. Google engineers, working in close collaboration with Google DeepMind, concluded that a single unified design could no longer satisfy both sets of demands with maximum efficiency.
TPU 8t: The Training Powerhouse
The TPU 8t is built specifically for high-throughput AI model training. Each superpod clusters 9,600 chips together, providing 121 exaflops of compute and two petabytes of shared memory connected via high-speed inter-chip interconnects (ICI). This configuration enables nearly linear scaling even for the most complex model architectures. According to Google's official announcement (published Wednesday, April 22, 2026, PDT), the chip delivers approximately 2.8× better price-to-performance than the previous seventh-generation Ironwood TPU, and can compress what previously required months of training time down to weeks. At its maximum deployment scale, Google says it can orchestrate more than one million TPU 8t chips in a single cluster using its Pathways and JAX software systems.
TPU 8i: Designed for Inference and Reasoning
The TPU 8i addresses the rapidly growing inference market, where deployed AI models must respond to millions of user requests simultaneously with low latency. A key architectural feature is its use of on-chip static random-access memory (SRAM): each TPU 8i contains 384 megabytes of SRAM, triple the amount found in the Ironwood generation. This large SRAM buffer allows the chip to keep model weights and intermediate computations close to the compute cores, dramatically reducing memory access bottlenecks. Google stated the TPU 8i achieves 80% better performance per dollar compared to its predecessor and is specifically architected to support Mixture of Experts (MoE) models and reinforcement learning (RL) tasks — capabilities that underpin next-generation AI reasoning systems.
AI Hypercomputer and Ecosystem Integration
Both chips slot into Google's AI Hypercomputer, a unified infrastructure stack that combines purpose-built hardware, open software frameworks, and flexible cloud consumption models. The platform also integrates new Axion-powered N4A CPU instances, which Google says deliver up to 30% better price-performance for agent runtime workloads compared to competing hyperscalers. The AI Hypercomputer is the same foundation that powers Google's own Gemini model family and enterprise AI services. Google has also announced ongoing collaboration with chip design partners Broadcom and Marvell Technology to expand its silicon supply chain, ensuring the infrastructure can scale to multi-gigawatt levels of capacity.
Growing Enterprise Adoption
Adoption of Google's TPU platform has expanded significantly ahead of this generation's launch. Citadel Securities has built quantitative research software on Google TPUs, and all 17 U.S. Department of Energy national laboratories use AI co-scientist software running on the platform. Anthropic has committed to consuming multiple gigawatts of TPU capacity from Google Cloud. Meta also signed a multiyear, multibillion-dollar agreement for TPU access, as reported by The Information in February 2026. These partnerships underscore the growing commercial momentum behind Google's custom silicon strategy as enterprises accelerate deployment of AI at scale.