Training Your
Own Models
Browse the book
Chapter 36 / 408 min read

36 Read MiMo GLM and other architecture families

The following are case studies in design choices. They are not a leaderboard or a claim that these are the newest or best models for every use. DeepSeek-V3 and Mamba-2 are deliberately included as clear reference designs. Newer names do not invalidate the mechanisms they illustrate.

Match the requested names to exact releases

Name in conversation Verified release or reference Why it is in this book
MiMo 2.6 XiaomiMiMo/MiMo-V2.6-Flash-RL and MiMo-V2.6-Pro-RL; newer Flash-MOPD and Pro-MOPD also listed officially Local/global attention, MoE, multimodal inputs, post-training and speculative decoding
GLM 5.3 Flash zai-org/GLM-5.3-Flash, with a separate BF16 release listed by Z.ai Linear plus sparse attention, MoE and mHC
Qwen larger hybrid Qwen/Qwen3.5-397B-A17B Gated DeltaNet plus attention, routed plus shared experts
DeepSeek reference deepseek-ai/DeepSeek-V3 MLA cache compression and an inspectable MoE inference implementation
Mamba reference Mamba-2 in state-spaces/mamba Structured recurrent state and hardware-aware scan algorithms
Liquid The exact small LFM checkpoints in the earlier practical chapter A small hybrid makes the sequence-mixer comparison accessible

The official MiMo Flash card says 309B total, the technical report rounds to 310B, and the Hub's tensor summary shows approximately 311B. These are different displayed counts; this book does not claim to have independently reconciled every auxiliary tensor or counting convention. It uses “about 310B” for rough storage calculations and preserves the exact source labels in the evidence ledger. A10 A11

The GLM card reports 320B total and 18B active; the MiMo cards report approximately 15B active for Flash and 42B for Pro. These are publisher specifications, not measured resident memory. A10 A12 A13

MiMo V2 6 as a local and global attention design

Start with the Flash checkpoint's configuration. The text backbone has width 4,096 and 48 layers. Its pattern contains 39 sliding-window layers and 9 global layers. The local window is 128 tokens. Global attention has 64 query heads and 4 KV heads; local attention has 64 query heads and 8 KV heads. Key/query width is 192 per head and value width is 128. There are 256 routed experts, with eight selected per token; the first FFN is dense. The exact context configuration is 1,048,576 positions. A14

These are compact architectural facts, not advice to use a million-token sequence in the earlier training script. You can now reason about them. Frequent narrow-window layers keep many attention operations local. Global layers provide direct access across the sequence. GQA reduces stored key/value heads. MoE increases stored FFN capacity without executing every expert for each token. The representation width need not equal query-head count multiplied by head width: projection matrices can expand the attention space and project back.

Pro scales the same general pattern: the official card lists 70 layers, including 60 local and 10 global, width 6,144, 384 routed experts and top-8 activation. Its global KV head count is eight. The different global-layer count and KV head count substantially change cache cost even before accounting for its roughly 1.02 trillion parameters. A12

Follow the input and output interfaces

The Flash release includes a vision configuration, audio configuration and special image, video and audio token identifiers. Its source defines vision patch embedding, a patch merger, an audio encoder and audio tokenization components alongside the text backbone. Those branches explain why a text-only tokenizer call is not a complete multimodal input pipeline. Inspect the processor's resizing, frame sampling and audio sample-rate assumptions with the checkpoint. A02 A14

The model card describes a vision encoder and two-stage audio input path, plus a separate five-layer speculative drafter. This book treats those as input-understanding and inference components rather than assuming native image or waveform output. A10

A configuration field is not the entire implementation

Three inspection details prevent expensive mistakes.

First, the released MiMoV2MoEGate.forward rejects training mode for its noaux_tc routing path. This is direct evidence that the inspected Hugging Face model implementation is not a ready-made full-training recipe. A separate research training framework can exist without making this particular loader trainable. A02

Second, the main configuration contains an MTP-related count, while the separate dflash/config.json describes a five-layer drafter and an eight-position block. Read the separate component rather than assuming one legacy field completely specifies speculative decoding. A15

Third, the quantization configuration combines quant_method=fp8 with store_dtype=mxfp4, and lists higher-precision exceptions. Storage format, runtime arithmetic and accumulator precision need not match. A filename or Hub dtype badge cannot tell you exactly how many bytes your chosen engine allocates. A14

Training data and post training implications

The MiMo technical report describes text-only pretraining followed by multimodal training. Reported totals are 48T tokens for Flash, split into 26T text-stage and 22T multimodal-stage tokens, and 30T for Pro, split into 27T and 3T. It then describes agent-focused mid-training, context extension, quantization-aware training and mixed-task reinforcement learning. The reported RL setup freezes the router and uses interactive trajectories with verifiers or graders. These are descriptions of that research run, not a public reconstruction of every original training example. A11

The practical lesson is about target information. A tool-use trajectory must include the relevant observations and consequences of an action; a code task needs a trustworthy test environment; a visual task needs the actual image. More compute cannot repair a reward that encourages the wrong behavior. The earlier tiny tool project should retain its safe local environment and held-out evaluator rather than reproduce broad autonomous web or security workloads.

The newer MOPD checkpoint is a post-training variant. Its card describes teacher-guided student continuations from several kinds of prefixes and a mitigation for repeated tool calls. It does not describe a new resident-parameter size that would make the model a consumer-card checkpoint. A distillation stage can preserve a student's size; “distilled” does not always mean “small.” A16

GLM 5 3 Flash as a recurrent and sparse attention design

GLM-5.3-Flash is a distinct newly trained base architecture, not merely a cheaper quantization of GLM-5.3. Its official card identifies a multimodal model, a hybrid of sparse and linear attention, mHC, and a 30T-token multimodal pretraining corpus. The broader GLM-5 technical report linked from the card is a family reference; it is not sufficient evidence for every Flash-specific implementation detail. A13 A17

The inspected configuration is more informative for this chapter. It specifies 45 text layers: 34 linear-attention layers and 11 sparse-attention layers. The first three FFNs are dense; the remaining 42 are sparse FFNs. They use 288 routed experts, eight selected experts and one shared expert. The hidden width is 4,096. The mHC stream multiplier is four, and the maximum configured position is 1,048,576. Its vision encoder is a separate 24-layer component. A18

The model's Transformers implementation names the linear mixer Kimi-style KDA. It has short convolution, a forget gate, a recurrent-state update for one-token decoding and a chunked path for sequences. The sparse-attention path has a separate indexer and compressed latent KV cache. This makes “hybrid” concrete: some layers carry a bounded recurrent state, while other layers retain growing historical information for selected attention. A19

Two shape changes to trace

Take a teaching input with B=1, T=16 and D=4,096. The initial text embeddings have shape [1, 16, 4096]. Four mHC streams give [1, 16, 4, 4096]. Before a sublayer, learned mixing collapses the stream axis; afterward, a write operation and constrained residual mixing update all streams. These are not four copies of the entire 320B-parameter model.

For the KDA layers, the config supplies 64 heads with width 128. A recurrent matrix with a 128×128 state per head has 64×128×128 entries per layer and sequence. At four bytes per entry, that is 4 MiB. Across 34 such layers, the illustrative matrix-state payload is 136 MiB for batch size one. This calculation excludes convolution buffers, sparse-layer caches, indexer state, temporary tensors and all weights. The inspected implementation explicitly stores the recurrent state in FP32. A18 A19

Sparse attention still needs a system

The indexer configuration includes a selection budget of 2,048 and grouping pools of four. That is an index-selection setting, not a 2,048-token context limit. The implementation handles incomplete tail pools and converts selections into a mask for the supported reference attention paths. Its reference path also expands compressed KV representations for the attention computation. Efficient production kernels can avoid costs that remain in such a readable implementation. A18 A19

The source supports cross-layer index reuse, but the inspected checkpoint's indexer_types are all full. Do not copy a feature advertised for another GLM version and assume it is enabled here. Likewise, the Transformers documentation explicitly says that its implementation omits the MTP layer, despite an MTP-related configuration field. Model support must be checked component by component. A18 A20

For training data, take the card's multimodal-corpus description as the available evidence. Do not invent exact image/text ratios, unpublished filtering thresholds or a full Flash pretraining script. A small architectural experiment can use controlled data to study recurrent versus sparse recall without claiming to reproduce the released model.

Qwen and DeepSeek make the design axes clearer

Qwen3.5-397B-A17B is an informative comparison because its official card specifies 15 repetitions of three Gated DeltaNet layers followed by one gated-attention layer. Its MoE has 512 routed experts, ten selected experts and one shared expert. The model has visual input capability and roughly 397B total versus 17B active parameters. Its documented native context is 262,144, with a separately described extension to approximately one million; the hosted product's settings should not be assumed to equal the local checkpoint's defaults. A21 A22

Compare the mechanism, not the scores. MiMo allocates many layers to short-window direct attention. Qwen allocates many to a recurrent linear mixer. GLM combines KDA with indexed sparse attention and adds multi-stream residual routing. All still need enough stored parameters and working memory for their chosen implementation. A similar active count does not imply identical compute, cache traffic, vision cost or latency.

DeepSeek-V3 is the older reference for the separation between MoE and latent attention: one changes which FFN parameters execute; the other changes how attention information is represented and cached. The original report describes 671B total and 37B active parameters, 14.8T pretraining tokens and an MTP objective. Its official inference code provides unusually direct places to inspect those mechanisms. It is included as a documented architectural reference, not as a claim about the latest DeepSeek release. A01 A04

Mamba-2 supplies the complementary reference: a language model can use structured state updates as its primary sequence mixer. An architecture family is a collection of implementation and modeling choices, not a single point on a “bigger is better” scale. The earlier small Liquid comparison lets you examine one hybrid on practical hardware before studying these much larger combinations. A06 A07