Training Your
Own Models
Browse the book
Chapter 25 / 404 min read

25 Add preferences and distillation only after supervised learning works

For this chapter's commands, open a fresh terminal at the companion root, run cd examples/llm, then source .venv-llm/bin/activate. If you skipped the first small-language-model project, complete its explicit environment setup first. Do not run these relative paths from an earlier project directory.

SFT gives the model examples of acceptable behavior. Sometimes you have two plausible answers and know which one is better. Sometimes a stronger model can produce useful training examples for a smaller one. These lead to preference optimization and distillation. They are extensions to a working supervised experiment, not substitutes for clear data and evaluation.

Preference data says which answer is better

A pairwise record contains the same prompt, a chosen response and a rejected response. Both responses should be judged under the same rubric. Here is an original example:

{
  "prompt": [{"role": "user", "content": "Stock lookup timed out. Tell the user."}],
  "chosen": [{"role": "assistant", "content": "The stock lookup timed out, so I don’t have a current count."}],
  "rejected": [{"role": "assistant", "content": "There are probably eight units available."}]
}

Do not create every rejected answer by appending obvious nonsense. Then the model can learn the easy surface distinction rather than the preference you care about. Include close comparisons: one answer quietly invents a deadline; another preserves uncertainty but is slightly longer.

Direct Preference Optimization, or DPO, compares the trainable policy to a fixed reference. Let x be the prompt, y+ the preferred completion and y- the rejected one. Define

z = beta × [(log p_theta(y+|x) - log p_ref(y+|x))

- (log p_theta(y-|x) - log p_ref(y-|x))].

The loss is -log sigmoid(z). Increasing z means the chosen answer gained relative to the rejected answer, measured against the reference. The original derivation connects this objective to a KL-regularized preference problem. beta is not a universal “more improvement” slider: it controls that reference-relative trade-off and scales the training signal. L29

TRL’s reviewed DPO interface accepts explicit prompt/chosen/rejected data. It also offers reference-log-probability precomputation. A frozen reference need not always occupy a second full trainable model’s memory, but its identity must be correct. For an SFT adapter followed by DPO, the reference should represent the intended SFT policy. Simply disabling an adapter can expose the original base instead, which is a different reference. L30

A reasonable first extension is short completions, the same sub-billion model, adapter updates and a small learning-rate experiment such as 5e-6 or 1e-5, with beta=0.1 as an initial comparison point. These are unmeasured starter values. Profile again: chosen and rejected sequences plus reference evaluation have a different memory and compute budget from SFT. Prefer off-line data before attempting an on-policy reinforcement-learning system that repeatedly generates, scores and updates.

Evaluate preference win rate and the earlier hard task metrics. A model can win a stylistic comparison while becoming less calibrated or less accurate. Track response length, unsupported claims and refusal/clarification behavior so preference learning cannot silently optimize the wrong shortcut.

Distillation transfers behavior rather than compressing a file

Knowledge distillation trains a student using information from a teacher. It is not weight quantization. Quantization changes representation precision; distillation changes what a usually smaller model learns. The original distillation framework uses teacher probability distributions, while sequence-level distillation can train on teacher-produced outputs. L31 L32

For this hardware budget, start with off-line sequence distillation:

  1. Select representative task prompts without using the final test set
  2. Generate candidate teacher responses using a method whose terms allow your intended use
  3. Validate facts, tool schemas and task success; reject bad examples
  4. Save accepted prompt/response pairs with teacher/version/provenance metadata
  5. Train the student with the same SFT pipeline used above
  6. Evaluate on untouched real cases and compare against human-authored SFT

This lets you run teacher generation and student training at different times. It avoids assuming both models fit in VRAM simultaneously. Teacher outputs are not automatically correct labels; fluent synthetic errors can be reproduced very efficiently by a student.

For logit distillation with matching token vocabularies, one common loss is

L_KD = tau² × KL(p_teacher^tau || p_student^tau).

Here tau is a temperature that softens the distributions and the squared factor compensates for temperature-related gradient scaling in the usual formulation. You often mix this with a supervised hard-label loss. If teacher and student use different tokenizers, their next-token distributions do not line up entry by entry. A naive KL across vocabulary positions is invalid. Sequence-level targets avoid that particular alignment problem.

The current TRL distillation documentation describes off-policy and on-policy behavior with a DistillationTrainer. The latter generates student trajectories and queries the teacher on those trajectories; it is a more demanding system than training on a saved JSONL file. Read the versioned API and memory path before adopting it. L33

A narrow student may outperform its own baseline on the target workload without reproducing the teacher’s broad reasoning ability. Distill the specific behavior you can test. For factual tools, distill the decision to consult evidence and the ability to use an observation, not a cache of the teacher’s guessed answers.