25 Add preferences and distillation only after supervised learning works
For this chapter's commands, open a fresh terminal at the companion
root, run cd examples/llm, then
source .venv-llm/bin/activate. If you skipped the first
small-language-model project, complete its explicit environment setup
first. Do not run these relative paths from an earlier project
directory.
SFT gives the model examples of acceptable behavior. Sometimes you have two plausible answers and know which one is better. Sometimes a stronger model can produce useful training examples for a smaller one. These lead to preference optimization and distillation. They are extensions to a working supervised experiment, not substitutes for clear data and evaluation.
Preference data says which answer is better
A pairwise record contains the same prompt, a chosen response and a rejected response. Both responses should be judged under the same rubric. Here is an original example:
{
"prompt": [{"role": "user", "content": "Stock lookup timed out. Tell the user."}],
"chosen": [{"role": "assistant", "content": "The stock lookup timed out, so I don’t have a current count."}],
"rejected": [{"role": "assistant", "content": "There are probably eight units available."}]
}
Do not create every rejected answer by appending obvious nonsense. Then the model can learn the easy surface distinction rather than the preference you care about. Include close comparisons: one answer quietly invents a deadline; another preserves uncertainty but is slightly longer.
Direct Preference Optimization, or DPO, compares the
trainable policy to a fixed reference. Let x be the prompt,
y+ the preferred completion and y- the
rejected one. Define
z = beta × [(log p_theta(y+|x) - log p_ref(y+|x))
- (log p_theta(y-|x) - log p_ref(y-|x))].
The loss is -log sigmoid(z). Increasing z
means the chosen answer gained relative to the rejected answer, measured
against the reference. The original derivation connects this objective
to a KL-regularized preference problem. beta is not a
universal “more improvement” slider: it controls that reference-relative
trade-off and scales the training signal. L29
TRL’s reviewed DPO interface accepts explicit prompt/chosen/rejected data. It also offers reference-log-probability precomputation. A frozen reference need not always occupy a second full trainable model’s memory, but its identity must be correct. For an SFT adapter followed by DPO, the reference should represent the intended SFT policy. Simply disabling an adapter can expose the original base instead, which is a different reference. L30
A reasonable first extension is short completions, the same
sub-billion model, adapter updates and a small learning-rate experiment
such as 5e-6 or 1e-5, with
beta=0.1 as an initial comparison point. These are
unmeasured starter values. Profile again: chosen and rejected sequences
plus reference evaluation have a different memory and compute budget
from SFT. Prefer off-line data before attempting an on-policy
reinforcement-learning system that repeatedly generates, scores and
updates.
Evaluate preference win rate and the earlier hard task metrics. A model can win a stylistic comparison while becoming less calibrated or less accurate. Track response length, unsupported claims and refusal/clarification behavior so preference learning cannot silently optimize the wrong shortcut.
Distillation transfers behavior rather than compressing a file
Knowledge distillation trains a student using information from a teacher. It is not weight quantization. Quantization changes representation precision; distillation changes what a usually smaller model learns. The original distillation framework uses teacher probability distributions, while sequence-level distillation can train on teacher-produced outputs. L31 L32
For this hardware budget, start with off-line sequence distillation:
- Select representative task prompts without using the final test set
- Generate candidate teacher responses using a method whose terms allow your intended use
- Validate facts, tool schemas and task success; reject bad examples
- Save accepted prompt/response pairs with teacher/version/provenance metadata
- Train the student with the same SFT pipeline used above
- Evaluate on untouched real cases and compare against human-authored SFT
This lets you run teacher generation and student training at different times. It avoids assuming both models fit in VRAM simultaneously. Teacher outputs are not automatically correct labels; fluent synthetic errors can be reproduced very efficiently by a student.
For logit distillation with matching token vocabularies, one common loss is
L_KD = tau² × KL(p_teacher^tau || p_student^tau).
Here tau is a temperature that softens the distributions
and the squared factor compensates for temperature-related gradient
scaling in the usual formulation. You often mix this with a supervised
hard-label loss. If teacher and student use different tokenizers, their
next-token distributions do not line up entry by entry. A naive KL
across vocabulary positions is invalid. Sequence-level targets avoid
that particular alignment problem.
The current TRL distillation documentation describes off-policy and
on-policy behavior with a DistillationTrainer. The latter
generates student trajectories and queries the teacher on those
trajectories; it is a more demanding system than training on a saved
JSONL file. Read the versioned API and memory path before adopting it.
L33
A narrow student may outperform its own baseline on the target workload without reproducing the teacher’s broad reasoning ability. Distill the specific behavior you can test. For factual tools, distill the decision to consult evidence and the ability to use an observation, not a cache of the teacher’s guessed answers.