AI Infra Interviews logo

The guide library

Understand the systems.
Build your way up.

Illustrated guides to AI infrastructure, from your first GPU to running a fleet. Browse the covers, read the samples and find your next step.

Every guide has a public preview. The accelerator guide is free with an account; the other PDFs are included with Premium.

All guides

5 illustrated PDFs

From one GPU to a fleet

The LLM Inference Systems Design Handbook

Follow a request through an LLM serving system. Work through cache capacity, batching, latency and scaling calculations, then practice the design decisions in mock interview rounds.

60 pages · 17 chapters · 8 mock rounds

Included with Premium

Run inference across a cluster

Distributed Inference on Kubernetes

Learn how a fleet routes requests, moves cached attention state and scales its workers. Compare NVIDIA Dynamo and llm-d through deployment examples, failure scenarios and capacity calculations.

96 pages · 25 chapters · 23 diagrams

Included with Premium

Choose and compare hardware

AI Accelerator Field Guide

Get familiar with the GPUs and TPUs used for AI. Compare memory and connectivity, estimate whether a model fits, and plan what to measure before choosing a machine.

47 pages · 37 hardware profiles · PDF

Free with an account

Understand the model you serve

Open Model Field Guide

Follow open model releases and the architecture changes behind them. Connect each design to its memory needs, deployment options and the checks that matter for your workload.

51 pages · 39 models · 8 labs

Included with Premium