Training Your
Own Models
Browse the book
Chapter 37 / 4013 min read

37 Connect architecture to data memory and deployment

A model's useful behavior depends on architecture, data, objective and execution together. The following design map turns the case studies into decisions for your own projects.

Choose data that tests the architectural claim

Design question Appropriate training or test evidence What a misleading test misses
Does local mixing suffice? Examples with controlled dependency distances, plus representative real sequences Short examples never exercise distant information
Does recurrent state retain useful details? Recall among distractors, multiple associations and changing sequence lengths One repeated pattern can fit in a simple summary
Does sparse attention select the right history? Evidence placed at varying positions, including competing plausible evidence A single obvious retrieval key is too easy
Does MoE capacity help? Balanced domain mixtures, held-out sources, routing and per-domain metrics Global loss can hide unused experts and domain regressions
Does a vision bridge learn grounding? Aligned images and questions, with image swaps and text-only controls Answers may come from textual shortcuts
Does audio input carry the answer? Aligned recordings and transcripts or descriptions, split by speaker/source as appropriate Recording duplicates and caption leakage inflate results
Does distillation transfer useful behavior? Teacher-generated examples with checks, then independent test tasks Student similarity to the teacher is not correctness
Does interactive RL improve actions? Safe environments, observable outcomes, robust verifiers and failure penalties An exploitable reward can improve while task success worsens

A tiny model can demonstrate any of these mechanisms. It will not establish the scaling behavior of a model trained on trillions of tokens. Hold one question fixed at a time: changing the tokenizer, data mixture, model size, training duration and attention type simultaneously prevents a clean architectural conclusion.

Pretraining and later objectives

Pretraining usually supplies broad prediction practice. For a causal language model, the target is the next token. A multimodal sequence model can condition the same text-prediction loss on image or audio features; additional modality objectives depend on the design. Continued pretraining or mid-training adds another broad or targeted phase before final task adaptation. These names describe stages, not universal algorithms.

Supervised fine-tuning uses desired responses or action traces. Preference optimization uses comparisons or preference signals. Reinforcement learning uses outcomes or reward estimates from generated behavior. Distillation lets a teacher provide targets or feedback for a student. An on-policy distillation setup trains on continuations produced by the student itself, so supervision covers mistakes the student is likely to make.

Multi-token prediction, or MTP, adds learning targets beyond the immediate next token. Speculative decoding uses a drafter to propose several tokens that a target model verifies. They are related in some designs but are not synonyms. A training auxiliary head might not ship in a runtime; a separate drafter might be trained after the target. Speed depends on proposal acceptance, the extra draft work, verification cost, batch size and kernel support. No fixed speedup follows from the words “predicts seven tokens.”

Data quantity is not the only scaling variable. Repeating duplicate documents increases processed-token count without proportionally increasing useful independent information. A much smaller model may be trained longer for economical inference, while another run may prioritize capability for a fixed training budget. Do not turn a single paper's token-to-parameter ratio into a universal rule for your narrow dataset.

Separate three kinds of sparsity and compression

MoE sparsity skips unselected experts for a token. Attention sparsity skips some query/key relationships. Weight sparsity sets or stores some individual weights as zeros. They require different algorithms and hardware support.

Unstructured pruning does not necessarily accelerate an ordinary dense matrix multiply: a stored zero still occupies its usual position and can still participate in the operation. Structured sparsity uses a supported pattern and representation. NVIDIA's cuSPARSELt documentation, for example, specifies particular 2:4 patterns for several dtypes, with other patterns for some formats. That is a kernel contract, not permission to assume any 50% sparse model runs twice as fast. A23

Quantization reduces precision. Distinguish weight storage, activation precision, arithmetic and accumulation, and cache precision. A four-bit weight file can still use higher-precision activations and accumulators. Scales, zero points, packing metadata and excluded tensors add overhead. Post-training quantization transforms a trained model; quantization-aware training exposes training to quantization effects. Quality checks must include the cases that already worried you, such as rare names, tool arguments and numerical values.

Distillation changes the training signal and can, when the student is smaller, change the model's resident size. Quantization changes the numerical representation of its parameters. Pruning changes which structure remains. None gives a free guarantee of preserved quality. Compare the exported result to your unchanged baseline on the same held-out cases.

Compute weight memory before discussing a GPU

Here are ideal payload calculations using rounded publisher parameter counts. A billion means 10^9, and a decimal GB means 10^9 bytes. These are not measured checkpoint sizes or usable VRAM requirements.

Study model Rounded resident count All BF16 payload All 8-bit payload Ideal packed 4-bit payload
MiMo V2.6 Flash 310B 620 GB 310 GB 155 GB
MiMo V2.6 Pro 1,020B 2,040 GB 1,020 GB 510 GB
GLM 5.3 Flash 320B 640 GB 320 GB 160 GB
Qwen3.5 397B A17B 397B 794 GB 397 GB 198.5 GB

The formula is parameter count × bits per parameter ÷ 8. Even the ideal four-bit columns exceed 24 GB several times over, before overhead or cache. The active count cannot replace the resident count in this equation. CPU offloading can move part of the storage to system RAM, but it also changes the execution system and its data-transfer cost. It does not turn the model into an all-resident 24 GB GPU model.

For full mixed-precision training, an illustrative Adam-style budget is two-byte working weights, two-byte gradients, four-byte master weights and two four-byte optimizer moments: 16 bytes per parameter before activations and temporary buffers. At 310B parameters that is 4.96 decimal TB. This is a stated-assumption illustration; actual optimizers and sharding schemes differ. At one billion it is already 16 GB, which is why the earlier projects carefully budget sequence length, microbatch size and optimizer state. At three billion it is 48 GB, so the selection ceiling is not a promise of straightforward full Adam training on one card.

Adapters reduce trainable state, not the requirement to make the frozen backbone available for forward and backward computations. The earlier QLoRA project works by combining a small base with quantized storage and low-rank updates. It should not be extrapolated to hundreds of billions of resident parameters simply because the adapter itself is small.

The cache can become a second large model sized allocation

Use MiMo's inspected dimensions to make a conditional cache estimate. Assume one sequence, BF16 keys and values, an efficient rolling cache for each local layer, and no speculative, vision, audio, padding or allocator overhead. For Flash, nine global layers each retain four KV heads of width 192+128. Thirty-nine local layers each retain up to 128 tokens with eight KV heads. A14

Flash cache bytes ≈ T × 9 × 4 × (192 + 128) × 2
                    + min(T,128) × 39 × 8 × (192 + 128) × 2

At 131,072 tokens this is approximately 2.84 GiB. At 1,048,576 tokens it is approximately 22.52 GiB. GiB uses 2^30 bytes. Pro's analogous calculation is approximately 6.29 GiB and 50.04 GiB, because it has ten global layers and eight global KV heads. These values are arithmetic predictions for the specified cache layout, not observations from a serving engine. A12 A14

Changing cache precision, sharing prefixes or using a different layout changes the result. A reference implementation that retains a full history for local layers can consume more. Batch size and simultaneous requests multiply working state. The important conclusion is robust: supporting a context length architecturally does not mean that it is free to serve.

In a hybrid recurrent model, recurrent state can remain fixed as T grows, but the attention layers' history still grows. GLM's 136 MiB recurrent-matrix illustration does not imply a 136 MiB total cache. Its latent attention cache, selection index state and buffers must also be counted. A careful estimator sums the actual state for every layer type rather than applying one formula to the entire model.

CPU and GPU inference have different bottlenecks

A CPU often offers more affordable addressable RAM and flexible control flow. A GPU usually offers much greater parallel arithmetic and memory bandwidth for supported operations. Neither statement alone determines latency for your model. The workload, precision, cache length, batch size, thread configuration and kernels matter.

At low-batch decoding, moving weights and cache entries can dominate. An illustrative bandwidth lower bound is bytes read per generated step divided by sustainable bandwidth. If an imaginary implementation really had to read 6 GB for a step and sustained 600 GB/s, memory traffic alone would take at least 0.01 seconds, before other work. This is not a benchmark for a named GPU. Caches, reuse and overlap can change the bytes that are actually read; contention and poor kernels can lower achieved bandwidth.

During prefill or larger batches, matrix multiplication can reuse weights across many tokens and become more compute-intensive. Multimodal input adds image/audio decoding and encoder work before text generation. Recurrent decoding may maintain little history but still execute substantial projection matrices. Sparse expert execution can save arithmetic while suffering from small per-expert batches or expensive data movement.

Measure at least time to first token, decode tokens per second, prompt tokens per second, peak memory and end-to-end task latency. Record input length, output length, batch/concurrency, precision and exact software revision. Separate cold-start loading from a warm request. Compare output quality at the same time. A faster server that silently truncates the useful context is not an equivalent result.

Apply the systems questions to each family

The following are engineering implications of the inspected structures, not measured rankings.

Case Data or evaluation emphasis CPU inference consideration GPU inference consideration
MiMo V2.6 Grounded multimodal examples and long tool trajectories; test local and distant evidence Large expert storage needs substantial RAM; offloaded expert traffic can dominate Local-window kernels help only part of the stack; global KV and expert residency remain
GLM 5.3 Flash Long-range recall plus multimodal grounding; test recurrent compression and sparse selection Runtime must support KDA state and sparse-index operations, with hundreds of GB of weight storage at common precisions Efficient recurrent and sparse kernels matter; count recurrent, latent and indexer state separately
Qwen3.5 397B A17B Visual grounding and tasks contrasting compressed history with direct attention A17B does not remove the resident 397B footprint Hybrid layers have different state and execution paths; shared experts always execute
DeepSeek V3 reference Text prediction and reasoning tasks; verify rare factual and structured outputs after compression Large expert memory and latent-attention operator support are both required MLA's cache advantage depends on the actual attention path and kernels
Mamba 2 reference Sequence tasks probing what fixed-size state retains and forgets Small recurrent state is useful, but projections and implementation quality still determine speed Efficient scan/chunk kernels help training and prefill; single-step decode is a different path
Small Liquid comparison The earlier narrow task plus distant-recall controls Measure the exact small checkpoint and CPU backend Check support for its convolution and attention blocks; do not infer speed from parameter count alone

Distributed memory introduces communication

Data parallelism runs model replicas on different examples. Ordinary replicas do not solve the problem of one replica being too large. Optimizer or parameter sharding, as in ZeRO-style or fully sharded approaches, divides selected training states across devices and gathers or communicates them when needed. Tensor parallelism splits large matrix operations across devices. Pipeline parallelism assigns successive layer groups to stages. Expert parallelism places different experts on different devices and dispatches token representations to them. Sequence or context parallelism distributes work or state along the sequence dimension; its exact meaning varies by framework. A24 A25

For an MoE model, expert dispatch may involve all-to-all communication. A lightly used remote expert can cost more to reach than its arithmetic suggests. Training adds backward communication and synchronization. Pipeline stages need enough microbatches to keep several stages busy. These mechanisms can be combined, but a count of “eight GPUs” omits the interconnect, memory per device, placement and communication schedule.

CPU/GPU offload is another placement strategy. A system with a 24 GB GPU and hundreds of gigabytes of system RAM is a different hardware proposition from a 24 GB card in a modest desktop. Feasibility depends on RAM, storage, transfer bandwidth, runtime support and acceptable speed. No one-card launch command is offered here for the giant study models.

Three bounded experiments that return you to the small projects

Experiment one checks routing and memory arithmetic

The companion examples/architecture_lab/inspect_designs.py uses only Python's standard library. It performs the original top-2 mixture example, computes the weight payload table and checks the cache calculations. It does not download a model, import a training framework or contact a service.

cd examples/architecture_lab
python inspect_designs.py

Expected evidence is a passed set of arithmetic assertions, the mixture output approximately [1.25, 1.50], and Flash's illustrative 1,048,576-token cache of approximately 22.52 GiB. These checks were run for the book. They verify arithmetic and a routing illustration, not model accuracy, VRAM allocation or generation speed.

Change k from two to one in a copy of the mixture example. Predict the new output before running it. Change batch size in the cache function and confirm proportional growth. Explain why reducing k does not change the full resident-weight payload. This is a useful architecture exercise before spending any GPU time.

Experiment two compares local and distant dependencies

Return to your already working miniature transformer. Keep its tokenizer, optimizer, train/test split and tiny parameter scale. As a proposed extension, compare a full causal mask with a local mask at window sizes 16 and 64. Use controlled sequences in which an early randomly generated key/value pair must be recalled after distractors. Hold out the random keys, values and templates, and vary the distance independently of sequence length.

Begin with a few hundred examples only to verify that the labels and masking are correct. Then use a fixed data and compute budget for the comparison. Include a task whose answer is near the end so that the local model has a fair positive control. Record exact-match recall by distance and held-out next-token loss. Inspect whether enough stacked local layers could transmit the information indirectly. Do not label a failure at one distance as a universal impossibility for every local-attention model.

This is an experiment plan, not a delivered and GPU-tested modification. It stays at the miniature project's parameter size, well below one billion. A dense attention implementation with a local Boolean mask may not save memory or time; use it first to study behavior, and study an efficient kernel separately.

Experiment three compares direct attention with compressed state

Use the earlier small Liquid comparison and a small dense baseline of similar practical size. Keep the real task and evaluation data fixed. Add tests for short local patterns, exact distant recall and a long sequence with several competing facts. Evaluate the unchanged checkpoints first; model size, pretraining data and tokenizer differences mean this is a product comparison, not a controlled proof about architecture.

If you need a causal architecture study, build two tiny models from the same starting design and change only the mixer, then train both from scratch on the same controlled dataset. Do not splice a recurrent block into a pretrained transformer and assume the remaining weights will preserve its behavior. A from-scratch architecture ablation and a comparison of pretrained model families answer different questions.

For either version, capture CPU and GPU measurements on the hardware you actually own. Do not extrapolate a published data-center throughput number to your desktop. Stop the study when you can explain the observed difference, reproduce the evaluation and identify which parts of the result are due to an uncontrolled difference. The goal is a defensible design decision, not a new state-of-the-art claim.

A final checklist for unfamiliar architectures

Before adding another family to your project, write a one-page inspection record.

  • Exact repository, revision, date, license and model class
  • Actual resident count and the publisher's definition of active count
  • Tokenizer and processor; input and output modalities
  • Layer types, attention heads, experts and residual structure
  • Per-layer persistent state and its precision
  • Full-training, adapter-training and inference support checked separately
  • Available data and objective evidence; unknown details clearly marked
  • Memory arithmetic with units and assumptions
  • A relevant quality baseline and a smoke-test plan
  • What you measured, what you calculated and what remains untested

Once you can complete this record, model names become much less mysterious. You can learn from a trillion-parameter architecture while making a practical choice to train a useful tiny model. The core skill is matching a mechanism and its data to a job you can evaluate and afford.