Training Your
Own Models
Browse the book
Chapter 20 / 405 min read

20 Teach domain conventions without turning weights into a database

For this chapter's commands, open a fresh terminal at the companion root, run cd examples/llm, then source .venv-llm/bin/activate. If you skipped the first small-language-model project, complete its explicit environment setup first. Do not run these relative paths from an earlier project directory.

The next project separates two goals often compressed into “teach it our business.” One goal is to recognize terminology and perform a workflow. The other is to answer exact questions about changing facts. They may need different solutions.

Use a fictional parts catalog so the experiment carries no confidential data. Part IDs look like P-101. A bin is a storage location. A stock count changes over time. You want the model to understand those conventions and to consult a lookup tool for the current count.

Start with the questions the model must answer

Build a held-out question set before choosing a training method. Include:

  • Terminology: “What does a bin identify?”
  • Document-grounded reasoning: “Using this paragraph, which field is missing?”
  • Current facts: “How many P-101 units are available now?”
  • Missing evidence: “What does P-101 cost?” when no price source exists
  • Conflicting versions: two manuals with different effective dates
  • Scope: a question about an unrelated catalog

A model may improve terminology while still hallucinating current counts. A single average score would hide that distinction. Track grounded correctness, unsupported assertions and abstention quality separately.

Retrieval-augmented generation, or RAG, retrieves relevant documents and supplies them as context at inference time. It can expose provenance and refresh information without changing model weights. The original RAG work combines learned generation with a non-parametric retrieval memory. Your implementation still needs retrieval evaluation, access controls and checks that the answer actually follows the retrieved text. L21

For exact inventory, prices, policy versions and permission-sensitive records, begin with retrieval or a database/tool lookup. For a consistent output format or specialist workflow, begin with SFT. For unfamiliar domain language throughout a substantial corpus, investigate continued pretraining. Combining these is often sensible: adapt how the model uses evidence while keeping evidence outside its weights.

Continued pretraining changes the text distribution

Continued pretraining, sometimes called domain-adaptive pretraining, starts from pretrained weights and continues the original language-modeling objective on new text. It does not require a question-answer pair for every paragraph. Research on domain/task-adaptive pretraining found benefits in the settings it studied; that does not guarantee improvement for every current decoder model or every corpus. L22

The provided cpt_train.jsonl uses records like:

{
  "source_id": "manual-train-01",
  "text": "Fictional Northstar catalog. Part IDs begin with P-. Stock counts are live data and must be looked up."
}

The code’s --task cpt option applies next-token loss to the text and an ending marker. It keeps documents separate to avoid hiding document-boundary questions inside a packing implementation. Use the proper base checkpoint for a real experiment, with its own verified revision, rather than casually treating an instruction checkpoint as interchangeable.

For a concrete smoke test with the base checkpoint, resolve its immutable revision first and record it:

BASE_REV="$(python -c 'from huggingface_hub import HfApi; print(HfApi().model_info("Qwen/Qwen3-0.6B-Base").sha)')"
python train_small_lm.py \
  --model Qwen/Qwen3-0.6B-Base --revision "$BASE_REV" \
  --task cpt --mode full --steps 3 --lr 5e-6 \
  --train data/cpt_train.jsonl --eval data/cpt_valid.jsonl \
  --out runs/domain-full-smoke

The architecture is still a causal decoder, but all document tokens are targets rather than only assistant answers. The saved full checkpoint is loaded directly for continuation. Give it a fresh prefix, without a chat template:

python generate_eval.py --run runs/domain-full-smoke \
  --device cuda --raw --max-new-tokens 64 \
  --prompt "Fictional Northstar manual. A bin identifies" \
  --output fresh-domain-gpu.jsonl
python generate_eval.py --run runs/domain-full-smoke \
  --device cpu --raw --max-new-tokens 64 \
  --prompt "Fictional Northstar manual. A bin identifies" \
  --output fresh-domain-cpu.jsonl

For an untouched base continuation, omit --run and supply --base-model Qwen/Qwen3-0.6B-Base --base-revision "$BASE_REV". CUDA uses BF16 inference and CPU FP32. The model, tokenizer, raw prefix and decode limit should otherwise match. A base continuation is not a reliable assistant response; converting it into an assistant requires appropriate instruction examples and a validated template.

A staged domain project might be:

  1. Clean and deduplicate licensed domain documents, preserving document and version metadata
  2. Hold out entire documents or organizations, not random adjacent chunks
  3. Measure domain and general-language baselines with the same tokenizer
  4. Run a short continued-pretraining experiment on the base model
  5. Add task-oriented SFT examples showing how to use the domain knowledge
  6. Compare against the original instruction model with retrieval

A useful experiment isolates effects. Compare base plus SFT against continued-pretraining plus the same SFT. Otherwise an improvement may come from the instruction data rather than the raw-text phase.

Watch catastrophic forgetting: gains on your narrow distribution can accompany regressions in general instruction following or other previously useful behaviors. Keep a small, fixed retention test set. A replay mixture of general examples is an experiment you can try, not a guarantee. Tune the mixture and learning rate using both target and retention metrics.

For a continued-pretraining run, an initial full-update rate such as 5e-6 on a sub-billion checkpoint is a cautious experiment, not a universal recommendation. A raw-text corpus with many tokens needs a token budget, held-out perplexity, checkpoints and a stop rule. Do not select an epoch count before estimating how many tokens an epoch contains.

Knowledge editing has a different promise

Ordinary fine-tuning can change factual behavior, but it is not a database update operation. The model does not expose a reliable row called “P-101 stock.” A repeated answer can become easier to generate while related questions remain wrong, old behavior persists under paraphrase, or unrelated answers drift.

Methods such as ROME and MEMIT study targeted changes to factual associations in particular model architectures and evaluation settings. Their existence demonstrates that factual behavior can sometimes be edited more selectively than with broad fine-tuning. It does not mean every fact has one isolated storage location or that edits are guaranteed to compose safely. L23 L24

Evaluate a proposed edit along four axes: does the requested answer change, does it generalize across relevant paraphrases, do unrelated facts remain stable, and do consequences of the edited fact stay consistent? Then test repeated edits and reversals. A success on one memorized prompt is weak evidence.

For this catalog, live stock belongs in an external store. Train the behavior “look it up and report the result.” Do not train tomorrow’s stock count into weights today. This choice also gives you a clearer audit trail and a straightforward way to correct errors.

Exercise. Put a changed stock count in the retrieved context while leaving training examples unchanged. Test whether the model follows the new evidence or repeats the old answer. Add a conflicting, older record with an explicit date and observe whether your evaluation catches version confusion.