Google TPU 8i
Google’s announced eighth-generation serving design. Larger SRAM and Boardfly target communication and state movement in reasoning workloads.
HBM
Published peak, not measured application throughput
The announcement distinguishes physical pod layout from active-chip count. Its Boardfly discussion includes up to 1,152 physical chips and 1,024 active chips; do not turn either into an assumed VM size.
When this is a sensible choice
Start here if…
Study 8i when all-to-all traffic and per-token synchronization are the concern. Its topology is designed differently from the training-oriented 8t, so version number alone is not the decision.
Choose another configuration if…
Before renting, verify availability, supported topology, checkpoint and runtime. The announcement’s communication improvements do not establish your model’s TTFT or TPOT.
Specifications with their boundaries attached
- Architecture
- TPU 8i
- Memory
- 288 GB HBM
- Memory bandwidth
- 8.6 TB/s per accelerator
- Peak compute
- 10.1 PFLOPS FP4 peak; other precisions not verified here
- Scale-up interconnect
- Boardfly ICI; Collectives Acceleration Engine
- Host attachment
- Managed TPU host; Boardfly
- Power
- Not published in checked per-chip reference
- Partitioning
- Slice/chiplet allocation; not NVIDIA MIG
- Catalogue status
- Announced; orderability not verified
Compute figures are theoretical peaks at the stated precision. Dense and structured-sparse rates must not be mixed. Bandwidth labelled bidirectional combines both directions. See the source documents.
Three bandwidths, three different jobs
Boardfly ICI; Collectives Acceleration Engine
DCN between slices
Where it appears in provider documentation
No provider configuration has been verified for this exact variant in this reference. Consult the linked manufacturer documentation; catalogue absence here is not evidence that the hardware is unavailable.
Model fit and software support
These publisher or serving-engine documents mention this hardware family. They have not been reproduced on our machines.
No model-specific recipe in our reviewed set certifies this exact hardware. The memory calculator can narrow candidates, but it cannot establish software support. Read the model register.
Tokens per second, TTFT and TPOT are not certified for this hardware in our reference. Use the benchmark checklist to compare an exact model, software revision and workload.
Nearby memory capacities, different tradeoffs
These are comparison candidates selected by memory capacity, not performance rankings or drop-in replacements.
Google TPU 8t
Google’s announced eighth-generation training design. It emphasizes large-scale pre-training, a 3D torus and a new scale-out network.
Google TPU7x (Ironwood)
Ironwood adds a much larger HBM budget and native FP8. Google’s technical name is TPU7x; its two chiplets are visible as two framework devices per chip.
Google TPU v5p
A large-pod training TPU with 95 GiB per chip, much more memory than v5e. The suffix identifies a different design, not a larger v5e VM.
Sources and review scope
Manufacturer and cloud documentation checked 2026-09-12. We reviewed specifications and the stated product boundaries; we did not run training or serving benchmarks on this device. Cloud memory and availability can differ by configuration.
- Google TPU 8t and 8i technical announcement ↗ · checked 2026-09-12
Cite this reference
AI Infra Interviews. Google TPU 8i: specifications and workload fit. Reviewed . https://aiinfrainterviews.com/hardware/tpu-8i
Include your access date when citing a changing specification. Link to the specification section for hardware figures or the explanation for a sizing or workload decision.
Research method and limits
We compile manufacturer specifications, cloud documentation, model cards and serving recipes. The reference preserves source units and distinguishes individual devices from nodes and racks. Conflicting figures and unknown fields remain labelled.
Our contribution is the comparison, unit reconciliation, worked arithmetic and workload explanation. Published peaks are vendor specifications. Calculator results are estimates under the displayed assumptions. Neither is a measurement from our own accelerator lab.
For a manufacturer’s specification, consult the original source documents. Cite our page when using its analysis, and retain primary-source attribution for underlying figures. This curated reference does not establish market share, live availability or a universal performance ranking.
Found a discrepancy? Send a correction with the page URL, exact variant, disputed figure and supporting primary document.
