17 Choose a model by its actual structure and files
After the first projects, model comparison becomes easier: you know which failure needs improvement and what your training budget contains. This section is a snapshot of official cards/configurations inspected on 2 October 2026. It does not claim that every listed checkpoint is the newest release in its family or that any is universally best.
A short list with the size traps made explicit
| Family and exact example | Published scale and structure | Consequence for this book |
|---|---|---|
Qwen/Qwen3-0.6B |
0.6B nominal; dense decoder; 28 layers; GQA with 16 query and 8 KV heads | Main full/adapter teaching checkpoint, Apache-2.0 L02 L34 |
Qwen/Qwen3-1.7B |
1.7B nominal; dense decoder; 28 layers; GQA | An adapter extension within the broader size range, not the full-update example L35 |
Qwen/Qwen3.5-0.8B and Qwen/Qwen3.5-2B |
Cards describe a vision encoder and a language model with alternating Gated DeltaNet and gated-attention layers; 24 LM layers | Newer hybrid/multimodal variants need different loading, module targeting and data handling; count all components, not only the LM label L36 L37 |
google/gemma-3-270m-it |
270M, with about 170M in embeddings and 100M in transformer blocks | Good small task-specialization candidate; large vocabulary makes “tiny blocks” different from “tiny total model” L38 |
google/gemma-3-1b-it |
Text-only small Gemma 3 variant; custom Gemma terms; 32K published context | Check exact unique count for a strict one-billion limit; use the correct Gemma template L39 L40 |
ibm-granite/granite-4.0-350m |
350M dense decoder; 28 attention layers; GQA; Apache-2.0 | Sub-billion candidate for a separate template-validated recipe L41 |
ibm-granite/granite-4.0-h-350m |
340M hybrid; attention plus Mamba2 layers | Different recurrence/kernel and adapter-target considerations, despite a similar name L42 |
ibm-granite/granite-4.0-1b |
The architecture table gives 1.6B, despite “1b” in the ID | Within the general range, outside the practical full-update ceiling L43 |
ibm-granite/granite-3.3-2b-instruct |
Dense decoder; config has 40 layers, width 2,048 and a 49,159-entry vocabulary | Older, still inspectable domain/tool candidate; do not assume an exact 2.000B count L44 L45 |
HuggingFaceTB/SmolLM2-360M-Instruct |
Small dense instruction checkpoint with published training resources | Another sub-billion baseline; use its own tokenizer/template L46 |
meta-llama/Llama-3.2-1B-Instruct |
Dense multilingual text model with GQA and a Llama community license | Useful ecosystem comparison; “1B” is rounded and the license differs from Apache L47 |
The main script intentionally accepts the reviewed Qwen3 text architecture, rather than pretending these rows are interchangeable. A new model needs a new verification pass: loader class, tokenizer, stop tokens, supported attention implementation, masks, trainable modules and parameter count.
The latest inspected Granite 4.2-3B card describes a dense reasoning model released in August 2026. Its nominal “3B” label is not proof of an exact ≤3,000,000,000 total. Treat rounded 3B-class models, including SmolLM3 and Llama 3.2 3B, as boundary comparisons until you count unique parameters; do not silently relax a strict size cap. L48 L49
Liquid models offer a small hybrid comparison
Liquid provides genuinely small models as well as larger ones. The older official LFM2 card gives exact counts of 354,483,968 for LFM2-350M, 742,489,344 for LFM2-700M, 1,170,340,608 for LFM2-1.2B and 2,569,272,320 for LFM2-2.6B. The first two fit the practical full-update size ceiling; the latter two belong to the broader adapter/comparison range. These checkpoints use the LFM Open License, so review those terms rather than assuming Apache or MIT. L55
The currently linked successor LiquidAI/LFM2.5-350M
retains 16 layers: ten double-gated short-convolution blocks and six
grouped-query-attention blocks, with a 65,536-entry vocabulary and a
32,768-token context. Its published card favors narrow extraction,
structured-output and tool-use workloads. Its vendor CPU speed numbers
are specific measurements under their setup, not predictions for your
machine. L56
A short convolution mixes a nearby window of sequence positions. An input-dependent gate controls how much information flows through that operation. The attention blocks provide another way to mix information across positions. This is a hybrid sequence model; “Liquid” does not mean that every released checkpoint is simply a classical continuous-time liquid neural network. The technical report describes architecture search that includes hardware constraints. L60
The causal text interface still supports next-token SFT and continued
pretraining. You do not need a new label type merely because some blocks
use convolutions. But kernel support, cache structure, padding behavior
and adapter placement need new checks. Transformers provides a native
Lfm2ForCausalLM interface with labels and cache arguments.
Liquid documents LoRA SFT through TRL, targeting attention projections
in its example. That establishes an available training path, not that
every convolution is adapted or that our Qwen-specific script can be
pointed at LFM unchanged. L57 L58
There is also a serialization difference: the inspected LFM2.5 card
describes Python-like calls inside its tool-call markers by default,
with JSON as a prompted alternative. A Qwen JSON-call parser is not
automatically the right parser. Never execute model-generated Python
with unrestricted eval; parse a restricted call structure
and validate it against an allowlist. This is an application boundary,
regardless of model family.
For a separate LFM experiment, begin with a text-only extraction
dataset, inspect its saved template, enumerate trainable modules and run
the same one-batch/save/reload checks. The official TRL guide currently
contains tokenizer= examples; the reviewed current trainer
uses processing_class=. Resolve that version difference
before a long run. For CPU deployment, the official guide provides
existing GGUF/llama.cpp paths. Adapting a model and exporting your
changed weights requires supported conversion and a fresh quality check;
an existing vendor GGUF is not your fine-tuned artifact. L57 L59
Total active and effective are not synonyms
A dense model normally uses all its dense transformer-block weights for each token’s forward computation. A mixture of experts, or MoE, routes a token through a subset of expert networks. This can reduce arithmetic per token relative to the total model capacity. It does not make unused experts cease to exist in memory.
Qwen3-30B-A3B is explicitly 30.5B total and 3.3B activated, with 128 experts and eight selected. It is outside this book’s three-billion-total scope. Even an ideal four-bit payload for 30.5B weights is about 15.25GB before metadata, non-quantized tensors, adapters and activations. That arithmetic is not a promise of trainability on a 24GB card. L50
Gemma 4 E2B has another naming convention: the card gives 2.3B effective parameters and 5.1B with embeddings, using per-layer embeddings. It is also outside a strict ≤3B total scope. Its card lists Apache-2.0, whereas the Gemma 3 examples above use Gemma terms. Do not assign a license to a checkpoint merely from its family name. L51
A checkpoint’s serialized tensor count can differ from the number of
independent trainable entries when input/output embeddings are tied.
Tied embeddings reuse the same matrix for looking up
token vectors and projecting hidden states toward vocabulary scores.
Qwen3-0.6B’s configuration sets tie_word_embeddings=true;
its inspected safetensors metadata reports about 0.752B stored entries
while its nominal model has about 0.6B unique parameters. Count the
instantiated model’s unique parameters and inspect actual storage rather
than inferring optimizer memory from a file-page badge. L03 L34
Read a configuration as an engineering document
Several fields matter immediately:
hidden_size: the width of each token’s internal vectornum_hidden_layers: how many transformation stages tokens pass throughintermediate_size: the feed-forward expansion widthnum_attention_heads: query-head countnum_key_value_heads: shared key/value-head count for GQAvocab_size: rows in token-related tables and width of vocabulary logitstie_word_embeddings: whether input/output tables share weightsmax_position_embeddingsand positional settings: implementation capacity, not your validated training length
For a dense linear map, parameter count is input width times output width, plus a bias if present. For an embedding table, it is vocabulary size times embedding width. These simple products explain why two “one-billion-ish” models can place capacity in different parts of the network and produce different memory peaks.
Grouped-query attention, or GQA, lets several query heads share fewer key/value heads. This reduces key/value storage relative to giving every query head its own keys and values. It does not make all training activations shrink by the same ratio. RoPE encodes positional relationships through rotations in attention coordinates. Sliding-window attention limits some layers to nearby tokens. Hybrid attention/state-space models combine different sequence-mixing mechanisms, so a standard full-attention memory formula is not a complete model for them.
Tokenizer choice is equally practical. Compare how many tokens your actual domain text uses, how identifiers and non-English terms split, whether special tokens are reserved and how the template encodes a tool call. A large vocabulary can shorten some sequences while increasing embedding and output-head costs. Never replace a pretrained tokenizer with another and assume the old embedding rows still have the correct meanings.