2 Apr 2026
gemma-4-26B-A4B-it
Sparse compute in a workstation-sized family
25.2B reported total · text + image + video
Gemma / Diffusion MoE
A different generation loop, not a drop-in speed setting

New to model sizes or GPU memory? Start with weights, parameters and quantization →
DiffusionGemma belongs in a research or latency-sensitive evaluation where you can tolerate a different serving path and assess final-answer quality carefully. Google recommends the autoregressive Gemma line for stronger production quality.
Publisher-reported, rounded model count; not an exact tensor census. Multimodal components and packaging can change the artifact size.
Evaluate it for
Experiments comparing completion time, refinement work and answer quality.
Choose another path when
A streaming API that assumes ordinary autoregressive token semantics or a quality requirement it has not met.
Release context
The June research release generates through iterative refinement rather than the standard one-token-at-a-time loop. That makes it historically useful even if your current service remains autoregressive.
Deployment starting points
Start with an exact artifact and an engine that supports it. A configuration below is evidence of a documented path; its device count does not promise a particular throughput or concurrency.
A streaming API that assumes ordinary autoregressive token semantics or a quality requirement it has not met.
The pinned model card establishes the architecture. It does not, by itself, prove that a particular GPU count runs this checkpoint at your target context. The capacity screen below helps rule out undersized allocations; a successful load and workload test are still required.
This is the index’s reported tensor payload, in decimal GB. It excludes file headers and runtime memory. Mixed precision, conversion and host/device placement determine how much GPU memory the loaded model needs.
Inspect the exact index metadata →This estimates inference memory for a hypothetical uniform precision. It is useful for rejecting an allocation that is too small. It does not establish a working deployment or estimate training memory.
A uniform-precision GPU count would hide this model’s component layout. Publisher-reported, rounded model count; not an exact tensor census. Multimodal components and packaging can change the artifact size. Use the documented artifact and placement path above; no GPU count is inferred here.
Architecture in practice
Measure final-answer quality and completion time.
Autoregressive TPOT is not a universal comparison.
Mechanism schematic based on the pinned model card and configuration. It explains a design principle; it is not a full implementation graph.
| Exact checkpoint | google/diffusiongemma-26B-A4B-it |
|---|---|
| Text model type | diffusion_gemma_text |
| Layers | 30 |
| Hidden width | 2,816 |
| Attention / KV heads | 16 / 8 |
| Head dimension | 256 |
| Routed / selected experts | 128 / Not verified |
| Configured positions | 262,144 |
| Documented extension | Not recorded for this checkpoint |
| Inputs → output | text, image, video → text |
| License metadata | apache-2.0 |
25 sliding attention layers; 5 full attention layers. These counts describe the configured pattern.
Open weights do not imply unrestricted use. Read the applicable license terms linked from the pinned card.
A useful comparison
These are the autoregressive quality and deployment baselines to keep in the experiment.
2 Apr 2026
Sparse compute in a workstation-sized family
25.2B reported total · text + image + video
2 Apr 2026
The dense control for Gemma’s MoE experiment
30.7B reported total · text + image + video
These are editorial comparison candidates. We have not run a matched quality or serving benchmark, so this is not a ranking.
Prepare to explain it
Explain why tokens per second alone can obscure the difference between refinement steps and accepted output.
Attention layouts · Quantization · Courses and worked examples
For a deployment evaluation, record exact weights, precision, engine, device count, interconnect, prompt/output lengths and concurrency. Report task success, errors, TTFT, TPOT and useful throughput together.
Our assessment is editorial judgment based on the linked architecture and deployment evidence. Model facts are publisher-reported or attributed to runtime maintainers; no independent GPU benchmark was run.
AI Infra Interviews, “diffusiongemma-26B-A4B-it: release, architecture and deployment”, checked 2026-09-13. Preserve this date and the exact checkpoint when citing.