AI Infra Interviews logo

September 2026 / Illustrated research edition

Open Model
Field Guide

What changed in open models, and what it means for the systems you build.

Follow the release timeline. Understand the architectural choices. Compare documented hardware paths and learn how to test the service you actually need.

The PDF and new editions are included while your Premium membership is active. Keep the copies you download. All 39 model cards and 8 lab profiles are public.

Open Model Field Guide, September 2026: editorial illustration of dense, sparse and separate-memory model architectures

Already a member? Sign in to download your copy →

Report sample: Read the model name before the specifications
Page 4Read the model name before the specificationsOpen image to read the sample ↗
Report sample: The recent releases, in order
Page 8The recent releases, in orderOpen image to read the sample ↗
Report sample: Kimi K3: the release is a new serving stack
Page 14Kimi K3: the release is a new serving stackOpen image to read the sample ↗
51illustrated pages
39models in release order
8lab profiles and timelines

A reading path, then a reference

Understand the changes before comparing the numbers.

  1. Start with the system

    Weights, architecture and runtime; local versus hosted deployment; why popularity and adoption are different signals.

  2. Follow the architectural changes

    Expert residency, conditional memory, sparse reads, K3’s hybrid state, V4.1’s two phases and alternative generation loops.

  3. Choose a deployment experiment

    Named hardware counts and precisions, a worked memory reversal, and checks for tool calling, schemas, latency and useful work.

  4. Read the labs and newest-first briefs

    Architectural lineages and chronological model briefs with source links, alternatives and routes back to the detailed live cards.

A page from the report explaining Kimi K3 hybrid state and its eight-GPU deployment configurations

Inside the report

A model release can change the whole serving stack.

Kimi K3 combines recurrent state and attention history. The report connects that design to cache reuse, a 2.8T expert store and the documented eight-GB300 or eight-MI350-class GPU configurations.

Read the full K3 assessment and sources →

The same approach carries through the guide: explain the mechanism, identify the remaining cost, and state what evidence would justify the deployment.

Evidence you can follow.

Every model brief links to its exact card, release evidence and live research notes. Hardware configurations distinguish engine-maintainer recipes from publisher memory claims. The diagrams explain mechanisms; the Nano Banana cover is conceptual artwork.

Sources checked 2026-09-13. Assessments are editorial judgments, not an independent quality leaderboard. No GPU benchmarks were run. Release notes, announcements and weight availability are labeled separately; unknown days are not invented.