Calculators · free · formulas shown
The numbers an AI infra loop expects you to do in your head
Four tools built on the same formula sheet the interview rounds use. Each one prints the calculation next to the result, because the goal is to leave able to do it on a whiteboard, not to depend on the page. Hardware specs are vendor dense peaks as of September 2026; model presets come from each model's published config.
KV cache & memory-fit calculatorCompute the KV cache per token, per sequence and per batch for Llama, Qwen, DeepSeek, Mixtral and gpt-oss, or a custom architecture, then check whether weights plus cache fit on a chosen GPU. Formula shown with every result.Open →GPUs and time-to-train calculatorEstimate training FLOPs with 6ND, wall-clock time on a fleet at a stated MFU, the GPU count for a deadline, the fleet cost, and the minimum GPUs needed to hold mixed-precision Adam state. Every assumption is visible.Open →Cost per million tokens calculatorDerive the cost of a million output tokens from GPU price, replica size, batch and context, using bandwidth-bound and compute-bound decode throughput estimates, or plug in your measured tokens per second.Open →Roofline calculatorDraw the roofline for H100, H200, B200, MI300X, TPU and Trainium, find the ridge point, and place a kernel by its arithmetic intensity to see whether it is memory-bound or compute-bound and how far below the roof it sits.Open →
The spec sheet behind the tools
Dense tensor peaks only. Vendor headline numbers usually include 2:1 structured sparsity, which doubles them and which no dense GEMM reaches. The ridge point is the peak divided by memory bandwidth, the intensity a kernel needs before compute becomes the limit.
| Accelerator | Memory | Bandwidth | bf16 dense | fp8 dense | Ridge (bf16) | Note |
|---|---|---|---|---|---|---|
| A100 80GB SXM | 80 GB | 2.04 TB/s | 312 TFLOPS | n/a | 153 FLOP/B | No FP8 tensor cores; INT8 is 624 TOPS dense. |
| H100 80GB SXM | 80 GB | 3.35 TB/s | 989 TFLOPS | 1,979 TFLOPS | 295 FLOP/B | |
| H200 141GB SXM | 141 GB | 4.8 TB/s | 989 TFLOPS | 1,979 TFLOPS | 206 FLOP/B | Same compute die as H100; the upgrade is memory capacity and bandwidth. |
| B200 (HGX) | 180 GB | 8 TB/s | 2,250 TFLOPS | 4,500 TFLOPS | 281 FLOP/B | Announced at 192 GB; the HGX B200 datasheet lists 180 GB usable. Numbers here are dense; NVIDIA's headline figures include 2:1 sparsity. |
| B300 (Blackwell Ultra) | 288 GB | 8 TB/s | 2,500 TFLOPS | 5,000 TFLOPS | 313 FLOP/B | Per-GPU figures derived from NVIDIA's GB300 NVL72 rack numbers divided by 72; FP64 was cut to about 1.2 TFLOPS, so this is an inference and low-precision part. Medium confidence on TDP. |
| MI300X | 192 GB | 5.3 TB/s | 1,307 TFLOPS | 2,615 TFLOPS | 247 FLOP/B | Measured LLM inference throughput has trailed the spec sheet more than on NVIDIA parts; treat the peak as a ceiling, not a forecast. |
| MI325X | 256 GB | 6 TB/s | 1,307 TFLOPS | 2,615 TFLOPS | 218 FLOP/B | Announced at 288 GB, shipping at 256 GB. |
| MI355X | 288 GB | 8 TB/s | 2,500 TFLOPS | 5,000 TFLOPS | 313 FLOP/B | CDNA 4, shipping since Q3 2025. AMD's peaks are dense (it does not market sparsity figures). Liquid-cooled; the MI350X is the 1,000 W air-cooled sibling at about 2.3 PFLOPS bf16. |
| TPU v5e | 16 GB | 0.82 TB/s | 197 TFLOPS | n/a | 240 FLOP/B | INT8 is 393 TOPS. Pods of 256 chips over ICI; sized for serving and small-model training. ICI is 400 GB/s bidirectional per chip. |
| TPU v6e (Trillium) | 32 GB | 1.64 TB/s | 918 TFLOPS | n/a | 560 FLOP/B | INT8 is 1,836 TOPS. 256-chip pods. ICI is 800 GB/s bidirectional per chip. |
| TPU v7 (Ironwood) | 192 GB | 7.37 TB/s | 2,307 TFLOPS | 4,614 TFLOPS | 313 FLOP/B | GA April 2026. First TPU with native FP8. Pods of 9,216 chips on a 3D torus; JAX and PyTorch only. |
| Trainium2 | 96 GB | 2.9 TB/s | 667 TFLOPS | 1,299 TFLOPS | 230 FLOP/B | Eight NeuronCore-v3 per chip; Trn2 UltraServers link 64 chips over NeuronLink. |
| Trainium3 | 155 GB | 4.9 TB/s | 671 TFLOPS | 2,517 TFLOPS | 137 FLOP/B | GA December 2025 on TSMC 3 nm. Memory is quoted as 144 GiB (about 155 GB decimal). Adds MXFP8/MXFP4 and structured sparsity; Gen2 UltraServers link 144 chips over NeuronLink v4. |
As of September 2026. Refreshed on each new generation; older parts stay because fleets keep them.
