35 Understand larger models using the parts you already know
You can understand a model that needs a data center without owning that data center. The earlier projects taught the pieces: a classifier turned features into scores; an embedding model learned representations; the tiny transformer mixed information across tokens; the image and audio projects turned signals into tensors; fine-tuning changed a useful behavior. Large architectures reorganize and scale these pieces.
This final part is an architectural study, after the practical projects. Its purpose is to help you read a model card, follow a configuration file, predict what data a design needs, and recognize the system cost hidden behind a model name. The hands-on training path remains centered on models below one billion parameters, with a broader adaptation and comparison range up to three billion unique parameters; any rounded boundary model must be identified explicitly. The hundreds-of-billions-parameter examples here are not one-card training recipes. No large checkpoint was downloaded or executed for this book.
The case studies were checked against primary model cards, configuration files and implementation code on 2 October 2026. They include the requested MiMo-V2.6 and GLM-5.3-Flash, plus selected reference designs that make their choices easier to understand. “Flash,” “active,” “hybrid” and “open” each need a precise explanation. None is a hardware budget by itself.
Six questions that make a model card readable
For any unfamiliar model, ask these questions in order.
- What enters and leaves? Text tokens, image patches, audio features or several kinds of input? Does the model output text, images, audio or action tokens? An audio input encoder does not imply a speech output decoder.
- What mixes positions? Full attention, local attention, recurrent state, convolution or a layer-by-layer mixture?
- What transforms each position? One feed-forward network or a routed selection of experts?
- What persists during inference? Weights, a growing key/value cache, a fixed-size recurrent state, a short convolution buffer or several of these?
- How was it taught? Prediction targets, demonstrations, preferences, teacher outputs, interactive rewards, and the data formats that make those losses meaningful?
- What did the release actually include? Weights, tokenizer, processor, inference code, training code, data descriptions, data itself and an applicable license are separate items.
The architecture is a program with learned numbers. The checkpoint is a particular set of those numbers. A training recipe explains how they were obtained. The serving system implements that program efficiently for requests. When one item is public, do not assume that the others are complete.
The feed forward network becomes a mixture of experts
In the tiny transformer, each token representation passes through the same feed-forward network, or FFN. This is a dense model in the usual language-model sense: the same major parameter blocks participate for every token. An embedding lookup still selects rows, so “dense” is a convention about the network, not a claim that every scalar is read on every operation.
A mixture of experts, or MoE, replaces some FFNs
with several FFNs and a learned router. An expert is a
neural subnetwork. It is not a separate chatbot, a human specialist or
necessarily an interpretable topic module. The router reads a token's
current representation, computes expert scores and chooses a small
number of experts for that token. The selected outputs are combined. The
next token can choose differently. DeepSeek's released inference
implementation makes this separation visible in its Gate,
Expert and MoE classes. A01
An original four-expert example is small enough to do by hand. Suppose the scores for one token are 0.50, 0.30, 0.15 and 0.05. With top-2 routing, choose experts 0 and 1. Normalize their selected scores by their sum, 0.80. Their mixing weights become 0.625 and 0.375. If their output vectors are [2, 0] and [0, 4], the result is:
0.625 × [2, 0] + 0.375 × [0, 4] = [1.25, 1.50]
┌─ expert 0 ── [2, 0] ── ×0.625 ─┐
token → router → top 2 ──┤ ├→ sum
└─ expert 1 ── [0, 4] ── ×0.375 ─┘
experts 2 and 3 exist, but do not process this token
This example starts with supplied scores; it does not implement the production router's learned logits, sigmoid, balancing corrections or grouped selection. Those choices differ across models. In MiMo's inspected code, a correction bias affects expert selection, while the combination weights come from the uncorrected selected scores. “Take softmax and select a few experts” is therefore an incomplete description of that checkpoint. A02
For a batch with B sequences, T tokens per sequence and representation width D, the input shape is [B, T, D]. Flattening the first two axes gives [B×T, D]. With E experts, routing scores have shape [B×T, E]. Selecting k experts gives expert indices and mixing weights of shape [B×T, k]. A dispatch operation gathers the appropriate token rows for each expert; a combine operation returns their weighted outputs to the original token positions.
Resident parameters are the weights that must be stored somewhere for the complete model to be available. Active parameters describe the subset used in a particular token's forward computation under the publisher's counting convention. Active count is useful for rough compute reasoning. Resident count determines the basic weight-storage requirement. Neither includes all cache, activation, workspace or communication costs.
Imagine eight experts with ten million parameters each and a ten-million-parameter shared backbone. Top-2 routing gives 90 million resident parameters and roughly 30 million active parameters for a token. Across many different tokens, every expert can be used. Keeping only the two most recently selected experts on a GPU requires fetching missing experts later; it does not delete the other sixty million expert parameters.
A shared expert is used for every token in addition to the routed experts. Some designs include one; MiMo-V2.6's routed FFNs do not. A shared expert can learn broadly useful transformations while routed experts supply extra conditional capacity. This is a design choice to test, not proof that the router has discovered neat human categories.
What changes during MoE training
The input data can still be ordinary next-token language-modeling sequences. It does not require a human-written “send this word to expert 7” label. The training problem gains an allocation problem: some experts may receive too many tokens while others receive too few. Load balancing encourages usable distribution. A capacity limit can bound how many token assignments an expert processes, but dropping or rerouting overflow changes the learning calculation. A release's actual method matters.
Log the expert assignment histogram, the fraction of overflowed assignments if applicable, and loss by data source. The global average loss can hide a collapsed router or an expert that rarely trains. A token-level top-k choice is discrete; a training implementation must specify how gradients, balancing signals and routing decisions interact. Reading an inference forward pass alone does not establish a correct training algorithm.
On a GPU, a small dense matrix operation can be easier to execute efficiently than many tiny expert operations. On a CPU, dispatch and irregular memory access also cost time. MoE improves a particular balance of capacity and computation; it does not guarantee that every small batch runs faster than a smaller dense model.
Attention choices solve different problems
Attention compares a query with available keys and uses the resulting weights to combine values. Queries and keys answer “which positions should influence this position?” Values supply the information to combine. The earlier transformer already did this with a causal mask so future tokens were unavailable.
Full attention and the growing cache
During a prompt's prefill, many token positions are processed together. In ordinary full causal attention, position t may attend to all previous positions and itself. The total number of allowed position pairs grows approximately with T squared. During decode, the model typically generates one new token per active sequence. A KV cache stores prior key and value vectors so that the entire prefix does not have to be projected again at every step. The new query still has to interact with the relevant cached history.
A cache is working memory for a particular request. It is not newly learned model knowledge. A longer cache consumes memory even when the weights never change. An application's stored conversation history, a search index and the neural KV cache are three different kinds of memory.
Multi-head attention has several query/key/value heads. Multi-query attention, or MQA, shares one key/value head across query heads. Grouped-query attention, or GQA, shares key/value heads within groups of query heads. Reducing KV heads reduces cached vectors and memory traffic; query heads can remain numerous. This sharing is part of the learned architecture, so arbitrarily changing a pretrained checkpoint's head counts is not a lossless runtime setting. A03
For a teaching configuration with 12 layers, 4 KV heads, 64 key coordinates, 64 value coordinates, 4,096 tokens, batch size 1 and two bytes per cached number:
cache bytes = batch × tokens × layers × KV heads
× (key width + value width) × bytes per number
= 1 × 4096 × 12 × 4 × (64 + 64) × 2
= 50,331,648 bytes = 48 MiB
Using eight KV heads would double this payload to 96 MiB. This is cache arithmetic, not total inference memory. Allocator overhead, temporary tensors and the weights still exist.
Sliding windows and selected positions
Sliding-window attention, or SWA, keeps each query's direct attention local to a recent window. With a fixed window W, local attention's position-pair count grows approximately with T×W instead of T squared. An efficient rolling cache can discard older keys and values for that local layer. Several local layers expand the effective receptive field, and occasional global layers allow direct long-range communication.
Sparse attention instead selects a subset of historical positions or blocks according to a pattern or learned selection mechanism. A learned indexer can look for useful distant positions that a fixed recent window would omit. It also introduces its own scores, state, computation and possible selection errors. Sparse attention is sparsity over token connections; MoE is sparsity over parameter blocks. A model can use both.
Sparse selection does not automatically create a constant-size cache. If a future query might select any past position, the system needs a way to retain, compress or recover the corresponding information. Masking most entries of a dense attention matrix is mathematically sparse, but can still allocate and compute like a dense implementation. Hardware-efficient sparsity requires a suitable kernel and cache layout.
Latent attention and exact kernels
Multi-head latent attention, or MLA, stores a
compressed representation from which the needed key/value information
can be obtained. Think of retaining a shorter learned code instead of a
separate full key and value for every head. DeepSeek-V3's official
inference code exposes both a direct path and an optimized absorbed
path; the latter maintains latent and positional caches. The choice of
algebra and kernel is part of whether the memory advantage appears in
practice. The kv_lora_rank configuration name here
describes a low-rank attention projection, not a separately trained user
LoRA adapter. A01 A04
FlashAttention addresses a different level of the system. It computes exact attention with a more memory-efficient execution order, tiling the work to reduce transfers between memory levels. It does not, merely by being enabled, replace full attention with a local window or a recurrent state, and it does not make the full-attention pair count linear. Numerical rounding can differ between kernels. A05
This distinction is useful when a model called “Flash” is discussed. A product suffix, FlashAttention, an FP8 checkpoint and a sparse attention pattern are four separate facts to verify.
Recurrent state and short convolution
A recurrent layer updates a state as tokens arrive. In a simplified notation, the new state is a function of the previous state and current token. The output reads from the new state. Its state can remain fixed in size as the sequence grows. That is a strong inference-memory advantage, but the state must summarize history rather than preserve arbitrary past token vectors separately.
A state-space model, or SSM, gives that state update a structured parameterization. Modern selective SSMs make parts of the update depend on the input. Mamba-2 is a useful reference because the original paper relates structured state-space computations to attention and explains efficient sequence algorithms. Its implementation contains convolution state, SSM state and separate sequence-processing and single-step paths. It is not simply a conventional transformer with a smaller KV cache. A06 A07
Linear attention refers to sequence-mixing formulations whose main sequence cost avoids the usual quadratic attention matrix. Some can be expressed as recurrent matrix updates. A schematic associative state is a matrix S that accumulates associations between key and value vectors; a query reads from S. A delta-rule-style update can correct an existing association rather than only add another one. Gating controls what is retained or forgotten. Kimi Delta Attention, or KDA, develops a finer-grained gated update in this family. A08
This is a conceptual analogy, not the exact KDA equation. Do not substitute a simple sum of key/value outer products for a production KDA implementation and call the result equivalent. Normalization, gates, delta corrections, chunking and numerical precision all matter.
A short convolution mixes a token with a fixed number of nearby positions using learned filters. A causal width-four convolution needs only a short recent buffer for incremental operation. A depthwise convolution applies a filter independently within each channel rather than combining every channel with every other one at that step. Other projections can still mix channels. Short convolution offers useful local processing; by itself it does not create unbounded long-range recall.
The earlier Liquid LFM comparison belongs here conceptually: its inspected small models mix gated short-convolution blocks with attention blocks. The brand name “Liquid” does not mean that every checkpoint uses the same continuous-time differential-equation architecture. Read the specific LFM version's block list. That small-model chapter supplies the practical comparison; there is no second large-model Liquid training recipe here.
Why hybrid models keep some attention
A hybrid model alternates or combines different sequence mixers. It may let recurrent or local layers do most of the repeated work while a smaller number of attention layers retain direct access to the past.
Example A local → local → local → global → repeat
Example B recurrent → recurrent → recurrent → sparse attention → repeat
Example C short convolution → attention → short convolution → …
These diagrams are patterns, not interchangeable checkpoint definitions. They explain the design tension: compressed history can be cheap; direct retrieval from history can be precise; both consume compute and memory in different ways. Evaluate a hybrid on exact copying, retrieval among distractors, long-range dependencies and the real downstream task. A short next-token loss alone cannot establish useful long-context behavior.
Multimodality connects encoders to a sequence model
A modality is a kind of input or output, such as text, images or audio. An encoder converts a signal into learned features. A projector maps those features to the representation width expected by the language backbone. A vision-language model can then process text embeddings and projected visual features in a common sequence.
text → tokenizer → token embeddings ──────────┐
image/video → patch encoder → projector ──────┼→ sequence backbone → output head
audio → audio encoder → projector ────────────┘
This original diagram describes a common input-fusion pattern. It does not imply that every family has all three branches or that one output head generates every modality. For image or audio generation, inspect the decoder, output representation and training objective separately.
A patch is a small image region or a group of nearby signal elements treated as one input unit. A 224×224 image divided into 16×16 patches gives 14×14 = 196 spatial patches before extra tokens or merging. A patch merger combines neighboring features to reduce the sequence length passed downstream. That reduces cost but also changes the detail the language backbone receives.
Video adds a time axis. Frame sampling, patch size, temporal grouping and image resizing all affect the token budget. Audio may be represented as spectrogram features, learned continuous embeddings or discrete codec tokens. A residual vector quantizer, or RVQ, uses several codebooks in sequence to represent a signal with discrete indices. Those indices are learned audio representations, not ordinary word-token IDs.
The correct data follows the desired cross-modal relationship. Image captioning needs image/text pairs. Document reading needs legible visual text and corresponding answers. Speech recognition needs audio with transcripts. Sound-event description needs audio with grounded descriptions. A pile of unpaired images and a pile of unrelated sentences does not, by itself, specify which sentence describes which image. Check alignment, time stamps, resolution, consent and splits by source recording or document.
Residual connections can have several streams
In the tiny transformer, a residual connection added a sublayer's output back to its input. Hyper-connections generalize the route through a deep network by maintaining several representation streams and learning how to read, write and mix them. Manifold-constrained hyper-connections, or mHC, constrain the stream-mixing operation to help control signal growth. The original work uses approximately doubly stochastic mixing: nonnegative entries with rows and columns normalized toward sums of one. A09
For intuition, two scalar streams [2, 10] mixed by rows [0.75, 0.25] and [0.25, 0.75] become [4, 8]. Their mean is preserved. That example explains constrained mixing; it does not prove the stability or usefulness of an entire trained network. The sublayers add learned information too, and finite normalization iterations introduce numerical approximation.
With four streams, a representation changes from [B, T, D] to [B, T, 4, D]. The sublayer can still operate on a collapsed [B, T, D] representation, then write back into the streams. More streams increase activation and memory-traffic costs even if the main expert width stays the same. GLM-5.3-Flash provides a concrete inspected implementation below.