A model’s name makes more sense when you can see what came before it. These profiles connect release history, architectural choices and useful cross-lab comparisons.
Kimi / 4 covered checkpoints
Follow the jump from a coding update to a new attention stack.
Moonshot’s K2-to-K3 sequence is useful because it contains two different kinds of progress. K2.5, K2.6 and K2.7-Code share the trillion-parameter class while changing the multimodal and coding workload emphasis. K3 changes the underlying deployment problem: a 2.8T expert store, hybrid attention and a newer kernel stack.
Latest covered: Kimi-K3 · 27 Jul 2026
Read the profile and release timeline →Qwen / 14 covered checkpoints
One family spans tiny dense models, coding MoEs and conditional memory.
Qwen offers the widest size ladder in this reference. That makes it useful for a disciplined evaluation: start with a small model, identify concrete failures, and move up only when a larger checkpoint fixes enough of them. The family also offers unusually clear dense-versus-MoE comparisons at similar stored size.
Latest covered: Qwen3.8-Flash-Next · 26 Aug 2026
Read the profile and release timeline →Gemma / 6 covered checkpoints
Gemma is a family of different design choices, not just different sizes.
Gemma 4 is a particularly useful learning family because its branches solve different problems. The April release includes effective-parameter E models, a conventional dense model and an MoE. June adds a unified encoder-free multimodal model and an experimental diffusion generator.
Latest covered: diffusiongemma-26B-A4B-it · 10 Jun 2026
Read the profile and release timeline →GLM / 6 covered checkpoints
Track the attention changes separately from the post-training gains.
GLM’s 2026 releases show why a version number is not an architecture description. The smaller GLM-4.7-Flash, the 744B-class GLM-5 line and the 320B GLM-5.3-Flash occupy different deployment classes. Within the flagship line, later training and attention changes deserve separate evaluation.
Latest covered: GLM-5.3-Flash · 26 Aug 2026
Read the profile and release timeline →DeepSeek V4 / V4.1 / 5 covered checkpoints
Flash evolves from a post-training update into a different memory system.
The April V4 preview, July Flash update, August Pro update and September V4.1 release form a useful chronological study. The July Flash update retains the earlier architecture and size; September changes the architecture. Those events should not be collapsed into one moving Flash label.
Latest covered: DeepSeek-V4.1-Flash · 10 Sept 2026
Read the profile and release timeline →Meta
Muse Glimmer / 1 covered checkpoints
A local-agent release includes the runtime around the model.
This reference currently covers Muse Glimmer, not the full historical Llama catalog or hosted Muse products. Glimmer is useful for understanding a local-agent release as a package: a dense multimodal model, quantized artifacts, a speculative drafter and device-specific evaluations.
Latest covered: Muse-Glimmer-30B · Aug 2026
Read the profile and release timeline →gpt-oss / 2 covered checkpoints
The older open release remains a useful compact reasoning baseline.
The August 2025 gpt-oss pair remains relevant because it exposes a practical small-versus-large reasoning comparison within one release. These are downloadable text models; they should not be described as the weights behind current hosted GPT services.
Latest covered: gpt-oss-120b · 5 Aug 2025
Read the profile and release timeline →MiniMax
MiniMax M3 / 1 covered checkpoints
Study how sparse attention changes what the GPU reads.
M3 is the MiniMax checkpoint covered in this edition. Its native multimodal training, million-position configuration and sparse-attention design make it a useful model for reasoning about long agent tasks. The June announcement and later open-weight availability are separate events.
Latest covered: MiniMax-M3 · 1 Jun 2026
Read the profile and release timeline →Profiles describe the models covered in this edition. They do not claim an exhaustive inventory or a ranking of lab quality.