DeepSeek-V4.1-Flash
The September model that changes the deployment diagram
Large multimodal services investigating conditional memory and distinct prefill/decode paths.
The open model reference / September 2026
You do not need to know what “35B-A3B” means yet. Start with the basics, then compare model releases, useful tasks and the hardware they need.
Explore by research lab






Compare all lab profiles. Coverage is the curated set in this edition, not each lab’s entire historical catalog.
New to open weights?
The layers and connections that define the computation.
The learned tensor values, sometimes converted or quantized.
Software that formats inputs, loads the model and generates outputs.
Hardware, request state, tools, access controls and evaluation.
You can run an open-weight model locally, operate a server, or call a provider that hosts it. Downloadable weights do not automatically include the materials needed to reproduce training, and local inference does not prevent an agent’s tools from sending data elsewhere.
Release timeline
Compare up to three checkpoints. A later date tells you where a model sits in the timeline; your workload decides whether it is the better choice.
39 of 39 checkpoints. Dates identify the sourced release or announcement event. Month-only records stay in their month; their position does not imply a day.
The September model that changes the deployment diagram
Large multimodal services investigating conditional memory and distinct prefill/decode paths.
Smaller than GLM-5.3, still a multi-GPU model
Visual coding agents and long-context multimodal services with compatible kernels.
The memory placement is part of the architecture
Multimodal services that can test GPU versus host placement and newer hybrid kernels.
A post-training release with a changed adoption decision
Controlled coding-agent and authorized defensive-security evaluations on a cluster.
A current dense baseline before you scale out
Mid-sized multimodal services and controlled comparisons with Qwen3.6-27B.
Keep the production update distinct from the preview
Reproducing the August Pro behavior and comparing it with the preview under one harness.
Large-scale text reasoning, with large-scale residency
Cluster-scale text reasoning and coding evaluation with a carefully pinned checkpoint.
A local agent needs more than a small download
Private on-device agents using text and image inputs on a capable workstation.
A post-training update you can compare fairly
Upgrades from V4 Flash preview and reproducible same-architecture comparisons.
A cluster model with a new attention design
Long repository tasks and multimodal agent evaluations on a high-memory cluster.
A million-token setting backed by index sharing
Large text-agent services investigating long-context retrieval and coding.
A coding upgrade without a new size class
Repository editing, test-driven repair and long agent runs.
A different generation loop, not a drop-in speed setting
Experiments comparing completion time, refinement work and answer quality.
One model path for text, vision and audio
Applications combining text, image, video and audio inputs.
Sparse reads do not automatically mean a small cache
Long-context coding and multimodal agent evaluation with supported MSA kernels.
The smaller V4 preview that made cache design central
Architecture study and controlled comparisons against the later Flash releases.
Trillion-parameter storage with sparse per-token work
Large-cluster text evaluations and study of expert-parallel serving.
A clean dense-generation comparison
Comparative multimodal evaluation under one serving configuration.
Keep the hardware fixed; compare the completed work
Existing K2.5 deployments evaluating agent and multimodal task improvements.
A smaller expert store for coding-agent experiments
Multimodal coding-agent prototypes and upgrades from Qwen3.5-35B-A3B.
Longer agent runs need process-level evaluation
Long coding and engineering workflows with recorded tool traces and acceptance tests.
Sparse compute in a workstation-sized family
Mid-sized text, image and video deployments and dense-versus-MoE studies.
The dense control for Gemma’s MoE experiment
Dense multimodal serving and comparisons that keep the model family fixed.
“Effective” is a compute label, not the download size
Small-device text, vision and audio experiments with a supported Gemma runtime.
More effective capacity, with tables still to place
Multimodal edge evaluation where E2B is too limited and the larger footprint is acceptable.
A cheap baseline for narrow tasks
Constrained text and visual tasks with task-specific acceptance tests.
A small step up when the tiniest model misses too much
Constrained text and visual tasks with task-specific acceptance tests.
A practical small-model starting point
Constrained text and visual tasks with task-specific acceptance tests.
Test whether the middle size earns its memory
Constrained text and visual tasks with task-specific acceptance tests.
A middle step before the largest Qwen3.5 model
Multi-GPU multimodal evaluation where the 35B model misses important cases.
Dense compute alongside a sparse sibling
Controlled dense-versus-MoE tests with text and visual inputs.
The easiest MoE memory lesson to make concrete
Learning MoE deployment and testing a comparatively modest multimodal service.
The launch model that established Qwen3.5’s hybrid design
Large-model baselines and study of the family’s initial architecture.
A large sparse-attention foundation for the GLM line
Large text-agent baselines and reproducing early GLM-5 results.
Executable feedback matters as much as model size
Code tools that can run tests, inspect repository state and measure completed changes.
The multimodal starting point for the K2 series
Reproducing K2-era evaluations and understanding large multimodal MoE deployment.
A compact way to learn latent-attention serving
Smaller text coding experiments and study of latent-attention caches.
A single large-memory GPU is a useful deployment class
Self-hosted text reasoning and controlled studies of scale versus the 20b model.
A compact reasoning baseline that still earns its place
Local text reasoning, tool-use prototypes and experiments with reasoning effort.
Architecture facts link to pinned model cards and configurations. Deployment rows name the artifact, precision, hardware count and runtime conditions. We label publisher memory claims separately from engine-maintainer recipes.
Our model assessments are editorial judgments, not an independent quality leaderboard. No GPU benchmark was run. Exact release days are not invented when only a month or announcement is documented. Download the public JSON records →