The Google Cloud Next keynote concluded recently had one big new announcement – The 8th-gen Tensor Processing Units (TPU) that beats the conventional one-size-fits-all chip and instead, split into 2 unique models of TPU 8t and TPU 8i.

Why you ask? because they want them to be specialized in one of the two big AI fields right now – massive training workload, and real-time inference, with the latter growing faster than ever due to the rise of agentic AI technology.

This slideshow requires JavaScript.

 

With models now expected to reason through problems, run multi-step workflows, and even learn from their own outputs, the infrastructure behind them needs to keep up. That’s why these new TPUs were developed alongside Google DeepMind, making sure they’re ready for the kind of workloads modern AI systems actually demand.

Specifically, TPU 8t excels at crunching through massive datasets faster, cutting model training time from months down to weeks. On the other side, TPU 8i is more like a reasoning engine, optimized for low-latency inference where multiple AI agents might be interacting with each other in real time, and even small delays can add up quickly.

On the other hand, TPU 8t is great at scaling – up to 9,600 chips in a single superpod for cutting-edge level compute with high efficiency, while TPU 8i focuses on limiting bottlenecks like memory limitations and latency, using things like larger on-chip memory and new acceleration engines to keep everything running smoothly.

Speaking of efficiency, power and cooling will always be the main concern alongside compute, therefore these new chips are said to provide up to 2x the performance-per-watt compared to the previous generation, augmented by liquid cooling and tighter integration.

Google TPU 8 Chip Announce (4)

Google is also continuing its long-standing approach of co-designing everything from silicon to software; these TPUs are tightly integrated with its own stack, including frameworks developers already use, while also being optimized for models like Gemini. At the same time, they’re keeping things open enough for broader adoption, with support for tools like PyTorch and JAX.

Both chips are expected to become generally available later this year as part of Google’s broader AI Hypercomputer platform, which bundles hardware, software, and orchestration into a single stack.

Facebook
Twitter
LinkedIn
Pinterest

Related Posts

Subscribe via Email

Enter your email address to subscribe to Tech-Critter and receive notifications of new posts by email.