18 Teach a small language model one useful behavior
A small generative language model turns a token sequence into a distribution over the next token. Around that simple interface, you can build quite different applications. Choose the application pattern before deciding to train.
| Pattern | What the model does | What the surrounding software must do |
|---|---|---|
| Completion | Continue a document or code prefix | Bound output, show suggestions, preserve the original |
| Chat assistant | Answer a structured conversation | Keep roles and history consistent with the model template |
| Structured extractor | Produce fields from supplied text | Validate a schema and reject unsupported values |
| Tool selector | Propose a function and arguments | Check permissions, validate arguments, execute safely, return observations |
| Retrieval assistant | Answer using supplied evidence | Retrieve authorized sources, track versions, check citations |
| Domain specialist | Interpret specialist language and workflows | Supply current facts, hold out realistic cases, retain general competence |
| Style rewriter | Express supplied content in a chosen voice | Preserve facts and commitments, review sensitive outputs |
These are system designs, not seven separate neural architectures. A single decoder model may support several after appropriate training. A good schema validator can make output structurally valid; it cannot make an incorrect field true. A tool-capable model proposes an action; the application decides whether to carry it out. Retrieved text is evidence, not permission to execute instructions found inside it.
Use prompting when the model already has the capability and needs a clearer task. Use retrieval or tools when it needs fresh, exact or permission-controlled facts. Use training when you need a repeatable learned behavior or genuinely new domain patterns and have representative examples. Begin with the narrowest pattern that solves your task, then add complexity only after a held-out test shows what is missing.
You have already built a small predictor and seen weights change during training. Now use the same ideas to customize a pretrained language model. The first project is deliberately narrow: turn a short list of facts into a concise update without adding a promise. You will begin with a small adapter, repeat the experiment with full fine-tuning, then extend the machinery to domain text and tool calls.
The practical full-update example uses Qwen3-0.6B, comfortably below one billion unique parameters. The broader model discussion stays around the requested three-billion-parameter ceiling and explicitly flags misleading size names. A model that can be loaded is not necessarily a model that can be trained. A training job that finishes is not necessarily a useful model.
Evidence labels. Model-card facts and library behavior have references. Memory arithmetic is a calculation with stated assumptions. Suggested learning rates, batch sizes and experiment stages are starter configurations, not measured optimums. The supplied Python files passed syntax checks and fourteen dependency-free unit tests. We did not install the training libraries, download model weights, run the tokenizer, train on a GPU, or measure model quality. The first task on your machine is therefore a smoke test, not a long training run.
The deliverable and the baseline
The deliverable is a local adapter directory, an exact base-model revision, a run manifest, and a comparison against the unchanged model. The target behavior is:
Input facts: “The export contains 47 items. Two images are missing. A retry time has not been agreed.”
Acceptable output: “The export contains 47 items, with two images missing. There’s no agreed retry time yet.”
Unacceptable output: “The export is complete. I’ll retry the two missing images tomorrow.”
The second answer sounds fluent, but changes completion status and invents a commitment. Your evaluator must notice that even if the prose is attractive.
Before training, try a clear instruction and two examples with the original checkpoint. Save those outputs. If prompting already meets the requirement, keep the simpler solution. This baseline is also how you discover what training needs to improve: length, factual preservation, formatting, or instruction following. These are different failure modes.
The included files live in examples/llm/.
make_data.py generated the supplied fictional fixtures.
train_small_lm.py handles training and masking.
generate_eval.py writes raw baseline or adapted
completions. score_tools.py handles a later structural
tool-call test. No example sends model weights, datasets, or training
logs to an external tracking service.
Make one record you can read
A supervised example pairs the situation the model sees with a
response you want it to produce. This is an original example in the same
shape as data/style_train.jsonl:
{
"source_id": "style-train-03",
"messages": [
{
"role": "system",
"content": "Write a concise update. Preserve the facts."
},
{
"role": "user",
"content": "The import processed 124 records. Three need a category."
},
{
"role": "assistant",
"content": "The import processed 124 records. Three still need a category."
}
]
}
JSONL means one complete JSON object per line. The pretty-printed
version above is for reading; the actual file uses one-line records.
source_id is your grouping key. All edits, paraphrases and
conversation turns from one original source stay in the same split.
Otherwise training can see an almost identical version of your supposed
test example.
The fixture contains only four training records. That is enough to test parsing and a few weight updates. It is not enough to learn a reliable writing style. Repeating four rows a thousand times does not create four thousand independent examples.
Why the template is part of the model
A chat model does not receive a Python list of roles directly. A
chat template converts the conversation into a token
sequence with role markers, turn boundaries and sometimes tool
definitions. The model learned the meaning of those markers during prior
training. Substituting another model’s markers is like changing the file
format while keeping the old parser. Transformers provides
apply_chat_template; use the template saved with the exact
checkpoint. L01
Qwen3-0.6B has a model-specific thinking switch. These examples
deliberately use enable_thinking=False and short outputs.
The training data does not require long reasoning traces. Its official
tokenizer template also handles tool calls and tool observations. The
checkpoint is pinned to commit
c1899de289a04d12100db370d81485cdf75e47ca, returned by the
official repository metadata when inspected. L02 L03
Open a fresh terminal at the companion root and prepare this project's isolated environment before invoking the trainer. Use Python 3.12. The CUDA 12.6 wheel below is an official build for a compatible NVIDIA driver and GPU; select another official supported build when required by your hardware. These pins were source-reviewed, not installed and runtime-tested during book preparation. F15
cd examples/llm
deactivate 2>/dev/null || true
python3.12 -m venv .venv-llm
source .venv-llm/bin/activate
python -m pip install torch==2.12.1 --index-url https://download.pytorch.org/whl/cu126
python -m pip install -r requirements-reviewed.txt
python -m pip check
python -m unittest test_offline.py
For a CPU-only inference environment, choose the official CPU wheel instead; QLoRA training and its CUDA quantization path are not a generic CPU training recipe. Save the actual resolved package list after your smoke test succeeds. Stay in examples/llm for the remainder of this chapter.
Run a preprocessing inspection before downloading weights:
python train_small_lm.py \
--train data/style_train.jsonl \
--eval data/style_valid.jsonl \
--out runs/inspect \
--inspect-only
With the environment above active, this command loads only configuration and tokenizer assets. Read both printed strings: the full input and the supervised target. The target should contain the intended assistant response and its ending marker. It should not contain the user’s instruction as something the model is being trained to answer with.
The code renders the prompt and the completed conversation separately. It verifies that the prompt is an exact prefix of the full text and of the token sequence. If this check fails, it stops. This is intentional. Some templates alter earlier reasoning content or serialize tools differently depending on later turns. A guessed boundary would quietly train the wrong tokens.
Loss masks teach the desired part
For this project, the prompt is context and the assistant response is
the target. Let the tokenized example be x[0], x[1], …, and
let m[t] be 1 for supervised response tokens and 0 for
prompt or padding tokens. The training loss is
L = - sum_t m[t] log p(x[t] | x[0:t]) / sum_t m[t].
In plain language: read the conversation, predict the next token, and average the surprise only over the tokens you want the model to learn to produce. The model still reads the prompt and uses it in its computation. Masking does not hide the prompt.
In the standard causal-language-model convention, labels use token
IDs at their own positions and the model shifts predictions and labels
internally. Do not shift a second time. Ignored labels are
-100, a sentinel used by the loss rather than a vocabulary
token. The collator pads input_ids, fills padding attention
positions with zero, and pads labels with -100. It masks by
position, never by “all tokens equal to the padding ID”; a model may
legitimately share its padding and ending token IDs. L04
You will meet three related settings in higher-level trainers:
- Completion-only loss: learn the completion after a prompt
- Assistant-only loss: learn assistant turns, excluding user, system and tool-observation text
- All-token language-modeling loss: learn ordinary text continuation, used later for continued pretraining
TRL’s reviewed SFT API exposes completion_only_loss,
assistant_only_loss and max_length.
Assistant-only masking depends on generation markers in the template or
a supported training-template patch. Do not assume every tokenizer
supports it. Our example makes the mask explicit so that you can inspect
it. L05 L06
Run the smallest adapter experiment
The environment README contains install commands and reviewed release
pins. They are not a fully tested lockfile. We inspected Transformers
5.18.0, Accelerate 1.15.0, PEFT 0.21.2, bitsandbytes 0.50.2 and TRL
1.14.1. The optional TRL path uses Datasets 5.0.1. PyTorch 2.12.1 with a
CUDA 12.6 wheel is an example supported by the official version page;
your driver and GPU must support the chosen wheel. Save
pip freeze, pip check, the CUDA runtime and
nvidia-smi output. L07 L08 L09 L10 L11 L12 L13
Begin with three optimizer updates:
CUDA_VISIBLE_DEVICES=0 python train_small_lm.py \
--train data/style_train.jsonl \
--eval data/style_valid.jsonl \
--out runs/style-lora-smoke \
--mode lora --steps 3
The checkpoint is a pretrained model: its original weights already encode useful language patterns. Fine-tuning starts from those weights and adapts behavior using your data. Supervised fine-tuning, abbreviated SFT, names the learning objective and data arrangement. LoRA names which parameters you update. You can perform SFT with LoRA or with all parameters trainable. These terms are not competing model types.
The first run uses these intentional starter choices:
| Setting | Value | Why it is there |
|---|---|---|
| Unique base size | about 0.6B | Small enough to compare methods rather than chase capacity |
| Microbatch | 1 example | Minimize simultaneous activation memory |
| Gradient accumulation | 8 microbatches | Combine small batches before an optimizer update |
| Maximum sequence length | 512 tokens | Bound the first experiment; over-length examples raise an error |
| Learning rate | 1e-4 |
Initial adapter experiment, to be checked on validation |
| LoRA rank | 8 | Small trainable correction |
| LoRA alpha | 16 | Scale multiplier alpha/r = 2 |
| LoRA dropout | 0.05 | Randomly suppress some adapter inputs during training |
| Gradient checkpointing | on | Recompute selected intermediates during backward |
| Attention implementation | SDPA | Use the framework implementation without a separate FlashAttention install |
| Optimizer | ordinary AdamW | Keep the optimizer accounting explicit |
| Weight decay | 0.01 | Set on TrainingArguments so Trainer builds the intended decay/no-decay groups |
| CPU offload | off | No hidden dependence on host-memory transfers |
The three-step run checks that loss is finite, parameters are trainable, a checkpoint saves, validation runs, and memory is logged. It does not establish that the model improved. A real experiment needs more independent examples, a fixed validation set, and a test set you do not use to tune settings.
LoRA is a learned correction matrix
Consider a linear transformation y = W x, where
W has d_out × d_in entries. A full update
changes those entries directly. LoRA freezes W and trains
two smaller matrices:
y = W x + (alpha/r) B A x,
with A of shape r × d_in and B
of shape d_out × r. The added trainable count is
r(d_in + d_out) rather than d_in × d_out. This
is the low-rank reparameterization introduced in the LoRA paper. L14
For a 1,024-by-1,024 layer and rank 8, the original layer has
1,048,576 entries. The adapter has 16,384. These are calculated values.
The adapter cannot represent every possible change to W at
this rank, but the useful change for your narrow task may fit within
that restriction.
Rank controls the capacity of this correction. Alpha changes its scale. Holding alpha fixed while changing rank changes both capacity and scale; do not treat such an experiment as “rank only.” The baseline uses alpha equal to twice rank so the standard scale stays at two. Other LoRA variants have different scaling rules, so save the complete adapter configuration.
target_modules="all-linear" asks PEFT to attach adapters
broadly to supported linear layers rather than just query and value
projections. Embedding tables and the output head are not automatically
equivalent to these targets. Print the actual trainable parameter names
and counts when you change architectures. A configuration copied from an
attention-only transformer can miss important modules in a hybrid model.
L15
Frozen weights still participate in forward computation, and gradients must propagate through the network to reach adapters. LoRA reduces the number of gradients and optimizer states you store. It does not eliminate activation memory or make training cost proportional only to the small adapter parameter count.
Inspect what the run produced
A successful run writes:
run_manifest.json: model ID, exact revision, arguments, installed package versions, GPU identity and dataset file hashesrun_results.json: before/after validation loss and measured peak allocated/reserved GPU memorycheckpoint-N/: a resumable training checkpoint at an optimizer stepfinal/: the deployment artifact and tokenizer
Resuming requires a complete checkpoint inside the original run directory. The script compares data hashes, model/revision, training configuration, package versions and recipe hash before writing anything. It preserves the original manifest and creates separate resume event/result files. Changing the planned step count changes the learning-rate schedule and is a new experiment.
In LoRA mode, final/ contains an adapter, not an
independent language model. It requires the same base architecture and
revision. In full-update mode, it contains the complete tuned model.
Keep optimizer/scheduler checkpoints for resuming training and
deployment artifacts for inference; they solve different problems.
Run the untouched baseline and the adapter on
style_test.jsonl. Read both outputs before looking at loss.
A decreasing validation loss is useful evidence about the probability of
reference answers. It is not proof of factual correctness, a specific
voice, or good decisions. Two equally good paraphrases can have
different reference likelihoods.
Exercise. Add one original training example that explicitly preserves uncertainty and one held-out case on a different topic. Compare the original and adapted model. Mark invented facts separately from stylistic improvements. Do not add the failed test case to training and keep calling it a test.