AI Infra Interviews logo

The open model reference / September 2026

Understand open models.
Choose what to run.

You do not need to know what “35B-A3B” means yet. Start with the basics, then compare model releases, useful tasks and the hardware they need.

Explore by research lab

See the lineage, not just the latest name.

Compare all lab profiles. Coverage is the curated set in this edition, not each lab’s entire historical catalog.

New to open weights?

The download is one part of the system.

01Architecture

The layers and connections that define the computation.

02Weights

The learned tensor values, sometimes converted or quantized.

03Tokenizer + runtime

Software that formats inputs, loads the model and generates outputs.

04Your service

Hardware, request state, tools, access controls and evaluation.

You can run an open-weight model locally, operate a server, or call a provider that hosts it. Downloadable weights do not automatically include the materials needed to reproduce training, and local inference does not prevent an agent’s tools from sending data elsewhere.

Learn parameters, quantization and your first deployment →

Release timeline

Newest first. Useful older baselines included.

Compare up to three checkpoints. A later date tells you where a model sits in the timeline; your workload decides whether it is the better choice.

39 of 39 checkpoints. Dates identify the sourced release or announcement event. Month-only records stay in their month; their position does not imply a day.

September 2026

Model release
DeepSeekMoE · text + image

DeepSeek-V4.1-Flash

The September model that changes the deployment diagram

Large multimodal services investigating conditional memory and distinct prefill/decode paths.

552Breported total

August 2026

Release note
Z.aiMoE · text + image + video

GLM-5.3-Flash

Smaller than GLM-5.3, still a multi-GPU model

Visual coding agents and long-context multimodal services with compatible kernels.

320Breported total18B active
Open-weight release
Alibaba QwenMoE · text + image + video

Qwen3.8-Flash-Next

The memory placement is part of the architecture

Multimodal services that can test GPU versus host placement and newer hybrid kernels.

180Breported total6B active
Announcement
Z.aiMoE · text

GLM-5.3

A post-training release with a changed adoption decision

Controlled coding-agent and authorized defensive-security evaluations on a cluster.

743Breported total39B active
Open-weight release
Alibaba QwenDense · text + image + video

Qwen3.8-27B

A current dense baseline before you scale out

Mid-sized multimodal services and controlled comparisons with Qwen3.6-27B.

27Breported total
General availability
DeepSeekMoE · text

DeepSeek-V4-Pro-0813

Keep the production update distinct from the preview

Reproducing the August Pro behavior and comparing it with the preview under one harness.

Not verifiedreported total
Open-weight release
Alibaba QwenMoE · text

Qwen3.8-2.4T-A95B

Large-scale text reasoning, with large-scale residency

Cluster-scale text reasoning and coding evaluation with a carefully pinned checkpoint.

2,400Breported total95B active
Model-card month
MetaDense · text + image

Muse-Glimmer-30B

A local agent needs more than a small download

Private on-device agents using text and image inputs on a capable workstation.

29.6Breported total

July 2026

Update release
DeepSeekMoE · text

DeepSeek-V4-Flash-0731

A post-training update you can compare fairly

Upgrades from V4 Flash preview and reproducible same-architecture comparisons.

284Breported total13B active
Open-weight release
Moonshot AIMoE · text + image + video

Kimi-K3

A cluster model with a new attention design

Long repository tasks and multimodal agent evaluations on a high-memory cluster.

2,800Breported total104B active

June 2026

Release note
Z.aiMoE · text

GLM-5.2

A million-token setting backed by index sharing

Large text-agent services investigating long-context retrieval and coding.

743Breported total39B active
Open-weight release
Moonshot AIMoE · text + image + video

Kimi-K2.7-Code

A coding upgrade without a new size class

Repository editing, test-driven repair and long agent runs.

1,000Breported total32B active
Research release
Google DeepMindDiffusion MoE · text + image + video

diffusiongemma-26B-A4B-it

A different generation loop, not a drop-in speed setting

Experiments comparing completion time, refinement work and answer quality.

25.2Breported total3.8B active
Model release
Google DeepMindDense · text + image + video + audio

gemma-4-12B-it

One model path for text, vision and audio

Applications combining text, image, video and audio inputs.

11.95Breported total
Announcement
MiniMaxMoE · text + image + video

MiniMax-M3

Sparse reads do not automatically mean a small cache

Long-context coding and multimodal agent evaluation with supported MSA kernels.

428Breported total23B active

April 2026

Preview release
DeepSeekMoE · text

DeepSeek-V4-Flash

The smaller V4 preview that made cache design central

Architecture study and controlled comparisons against the later Flash releases.

284Breported total13B active
Preview release
DeepSeekMoE · text

DeepSeek-V4-Pro

Trillion-parameter storage with sparse per-token work

Large-cluster text evaluations and study of expert-parallel serving.

1,600Breported total49B active
Open-weight release
Alibaba QwenDense · text + image + video

Qwen3.6-27B

A clean dense-generation comparison

Comparative multimodal evaluation under one serving configuration.

27Breported total
Model release
Moonshot AIMoE · text + image + video

Kimi-K2.6

Keep the hardware fixed; compare the completed work

Existing K2.5 deployments evaluating agent and multimodal task improvements.

1,000Breported total32B active
Open-weight release
Alibaba QwenMoE · text + image + video

Qwen3.6-35B-A3B

A smaller expert store for coding-agent experiments

Multimodal coding-agent prototypes and upgrades from Qwen3.5-35B-A3B.

35Breported total3B active
Release note
Z.aiMoE · text

GLM-5.1

Longer agent runs need process-level evaluation

Long coding and engineering workflows with recorded tool traces and acceptance tests.

744Breported total40B active
Open-weight release
Google DeepMindMoE · text + image + video

gemma-4-26B-A4B-it

Sparse compute in a workstation-sized family

Mid-sized text, image and video deployments and dense-versus-MoE studies.

25.2Breported total3.8B active
Open-weight release
Google DeepMindDense · text + image + video

gemma-4-31B-it

The dense control for Gemma’s MoE experiment

Dense multimodal serving and comparisons that keep the model family fixed.

30.7Breported total
Open-weight release
Google DeepMindDense · text + image + video + audio

gemma-4-E2B-it

“Effective” is a compute label, not the download size

Small-device text, vision and audio experiments with a supported Gemma runtime.

5.1Breported total
Open-weight release
Google DeepMindDense · text + image + video + audio

gemma-4-E4B-it

More effective capacity, with tables still to place

Multimodal edge evaluation where E2B is too limited and the larger footprint is acceptable.

8Breported total

March 2026

Open-weight release
Alibaba QwenDense · text + image + video

Qwen3.5-0.8B

A cheap baseline for narrow tasks

Constrained text and visual tasks with task-specific acceptance tests.

0.8Breported total
Open-weight release
Alibaba QwenDense · text + image + video

Qwen3.5-2B

A small step up when the tiniest model misses too much

Constrained text and visual tasks with task-specific acceptance tests.

2Breported total
Open-weight release
Alibaba QwenDense · text + image + video

Qwen3.5-4B

A practical small-model starting point

Constrained text and visual tasks with task-specific acceptance tests.

4Breported total
Open-weight release
Alibaba QwenDense · text + image + video

Qwen3.5-9B

Test whether the middle size earns its memory

Constrained text and visual tasks with task-specific acceptance tests.

9Breported total

February 2026

Open-weight release
Alibaba QwenMoE · text + image + video

Qwen3.5-122B-A10B

A middle step before the largest Qwen3.5 model

Multi-GPU multimodal evaluation where the 35B model misses important cases.

122Breported total10B active
Open-weight release
Alibaba QwenDense · text + image + video

Qwen3.5-27B

Dense compute alongside a sparse sibling

Controlled dense-versus-MoE tests with text and visual inputs.

27Breported total
Open-weight release
Alibaba QwenMoE · text + image + video

Qwen3.5-35B-A3B

The easiest MoE memory lesson to make concrete

Learning MoE deployment and testing a comparatively modest multimodal service.

35Breported total3B active
Open-weight release
Alibaba QwenMoE · text + image + video

Qwen3.5-397B-A17B

The launch model that established Qwen3.5’s hybrid design

Large-model baselines and study of the family’s initial architecture.

397Breported total17B active
Release note
Z.aiMoE · text

GLM-5

A large sparse-attention foundation for the GLM line

Large text-agent baselines and reproducing early GLM-5 results.

744Breported total40B active
Announcement
Alibaba QwenMoE · text

Qwen3-Coder-Next

Executable feedback matters as much as model size

Code tools that can run tests, inspect repository state and measure completed changes.

80Breported total3B active

January 2026

Model release
Moonshot AIMoE · text + image + video

Kimi-K2.5

The multimodal starting point for the K2 series

Reproducing K2-era evaluations and understanding large multimodal MoE deployment.

1,000Breported total32B active
Release note
Z.aiMoE · text

GLM-4.7-Flash

A compact way to learn latent-attention serving

Smaller text coding experiments and study of latent-attention caches.

30Breported total3B active

August 2025

Open-weight release
OpenAIMoE · text

gpt-oss-120b

A single large-memory GPU is a useful deployment class

Self-hosted text reasoning and controlled studies of scale versus the 20b model.

117Breported total5.1B active
Open-weight release
OpenAIMoE · text

gpt-oss-20b

A compact reasoning baseline that still earns its place

Local text reasoning, tool-use prototypes and experiments with reasoning effort.

21Breported total3.6B active

What counts as evidence here?

Architecture facts link to pinned model cards and configurations. Deployment rows name the artifact, precision, hardware count and runtime conditions. We label publisher memory claims separately from engine-maintainer recipes.

Our model assessments are editorial judgments, not an independent quality leaderboard. No GPU benchmark was run. Exact release days are not invented when only a month or announcement is documented. Download the public JSON records →