AI Infra Interviews logo
← Model reference

8 research labs / September 2026

Understand the lab.
Follow the models.

A model’s name makes more sense when you can see what came before it. These profiles connect release history, architectural choices and useful cross-lab comparisons.

Kimi / 4 covered checkpoints

Moonshot AI

Follow the jump from a coding update to a new attention stack.

Moonshot’s K2-to-K3 sequence is useful because it contains two different kinds of progress. K2.5, K2.6 and K2.7-Code share the trillion-parameter class while changing the multimodal and coding workload emphasis. K3 changes the underlying deployment problem: a 2.8T expert store, hybrid attention and a newer kernel stack.

Latest covered: Kimi-K3 · 27 Jul 2026

Read the profile and release timeline →

Qwen / 14 covered checkpoints

Alibaba Qwen

One family spans tiny dense models, coding MoEs and conditional memory.

Qwen offers the widest size ladder in this reference. That makes it useful for a disciplined evaluation: start with a small model, identify concrete failures, and move up only when a larger checkpoint fixes enough of them. The family also offers unusually clear dense-versus-MoE comparisons at similar stored size.

Latest covered: Qwen3.8-Flash-Next · 26 Aug 2026

Read the profile and release timeline →

Gemma / 6 covered checkpoints

Google DeepMind

Gemma is a family of different design choices, not just different sizes.

Gemma 4 is a particularly useful learning family because its branches solve different problems. The April release includes effective-parameter E models, a conventional dense model and an MoE. June adds a unified encoder-free multimodal model and an experimental diffusion generator.

Latest covered: diffusiongemma-26B-A4B-it · 10 Jun 2026

Read the profile and release timeline →

GLM / 6 covered checkpoints

Z.ai

Track the attention changes separately from the post-training gains.

GLM’s 2026 releases show why a version number is not an architecture description. The smaller GLM-4.7-Flash, the 744B-class GLM-5 line and the 320B GLM-5.3-Flash occupy different deployment classes. Within the flagship line, later training and attention changes deserve separate evaluation.

Latest covered: GLM-5.3-Flash · 26 Aug 2026

Read the profile and release timeline →

DeepSeek V4 / V4.1 / 5 covered checkpoints

DeepSeek

Flash evolves from a post-training update into a different memory system.

The April V4 preview, July Flash update, August Pro update and September V4.1 release form a useful chronological study. The July Flash update retains the earlier architecture and size; September changes the architecture. Those events should not be collapsed into one moving Flash label.

Latest covered: DeepSeek-V4.1-Flash · 10 Sept 2026

Read the profile and release timeline →

Muse Glimmer / 1 covered checkpoints

Meta

A local-agent release includes the runtime around the model.

This reference currently covers Muse Glimmer, not the full historical Llama catalog or hosted Muse products. Glimmer is useful for understanding a local-agent release as a package: a dense multimodal model, quantized artifacts, a speculative drafter and device-specific evaluations.

Latest covered: Muse-Glimmer-30B · Aug 2026

Read the profile and release timeline →

gpt-oss / 2 covered checkpoints

OpenAI

The older open release remains a useful compact reasoning baseline.

The August 2025 gpt-oss pair remains relevant because it exposes a practical small-versus-large reasoning comparison within one release. These are downloadable text models; they should not be described as the weights behind current hosted GPT services.

Latest covered: gpt-oss-120b · 5 Aug 2025

Read the profile and release timeline →

MiniMax M3 / 1 covered checkpoints

MiniMax

Study how sparse attention changes what the GPU reads.

M3 is the MiniMax checkpoint covered in this edition. Its native multimodal training, million-position configuration and sparse-attention design make it a useful model for reasoning about long agent tasks. The June announcement and later open-weight availability are separate events.

Latest covered: MiniMax-M3 · 1 Jun 2026

Read the profile and release timeline →

Profiles describe the models covered in this edition. They do not claim an exhaustive inventory or a ranking of lab quality.