Training Your
Own Models
Browse the book
Chapter 39 / 408 min read

39 Exercise checkpoints and worked answers

Use answers to check reasoning rather than memorize settings

Some exercises have a numerical answer. Others ask you to design an experiment, label real data, or make a tradeoff. Those do not have one universal correct configuration. The checkpoints below show what a defensible answer contains. A result can differ from an example and still be sound if its assumptions and evidence are explicit.

The first classifier and simple numbers

For predicted packing time = 0.7 times item count plus 2, five items give 5.5 minutes. Increasing the bias by one adds one minute to every prediction. Increasing the weight from 0.7 to 0.8 adds half a minute for five items and one minute for ten. This demonstrates the difference between an additive offset and an input-dependent effect.

The first two-feature classifier has two weights and one bias. Its architecture can represent a straight-line boundary. An XOR-style target requires a nonlinear representation or model. More gradient-descent steps do not turn a linear score into an arbitrary nonlinear boundary. A useful answer to the XOR experiment distinguishes model capacity from optimization success.

A valid explanation of its saved artifact names the feature order, three parameters, sigmoid, threshold, and target meaning. Saving only the weights while forgetting feature order would leave an incomplete inference contract.

Shapes and parameter counts

Sixteen sequences of 128 token IDs have shape [16, 128]. Replacing each ID with a 64-dimensional vector gives [16, 128, 64]. Token IDs are integer indices; the embedding values are floating-point learned representations.

A 256-row embedding table of width 128 has 32,768 parameters. A 50,000-row table at the same width has 6,400,000. Vocabulary size therefore matters greatly in a small language model. Tying the input and output table avoids storing a second independent copy of that same-shaped projection; it does not eliminate the computation of output logits.

An RGB batch of eight images at 96 by 96 pixels commonly has shape [8, 3, 96, 96]. A three-class classifier returns logits [8, 3]. A one-object detector may return objectness [8] and boxes [8, 4]. A binary segmenter may return pixel logits [8, 1, 96, 96]. These output shapes imply different targets and losses.

Loss weighting and accumulation

Microbatch size two, accumulation eight, and 300 optimizer steps produce 2 times 8 times 300 = 4,800 example presentations if all microbatches are full. The number of unique examples depends on sampling and repetition. If each example contains 128 supervised targets, the same schedule presents 614,400 supervised targets.

Suppose microbatch A has 100 supervised tokens with mean loss 2.0 and microbatch B has 1,000 with mean loss 1.0. Averaging the two means equally gives 1.5. Averaging all token losses gives (100 times 2 + 1,000 times 1) / 1,100, approximately 1.091. Neither calculation is mysterious, but they optimize different weighting. State whether your objective treats examples, microbatches, or supervised tokens equally.

In assistant-only training, a prompt token can be visible through attention while excluded from loss. The attention mask controls information flow; the loss mask controls which predictions are rewarded. A correct answer mentions both separately and verifies the actual supervised span after tokenization.

Memory and time arithmetic

Under the explicit FP32 AdamW assumption of sixteen persistent bytes per trainable parameter, 250 million parameters require 4 billion bytes, or about 3.73 GiB, before activations and other overhead. One billion requires 16 GB, about 14.90 GiB. Three billion requires 48 GB, about 44.70 GiB. These are inventories, not observed GPU peaks.

A dataset of 40,000 examples averaging 500 supervised tokens contains about 20 million supervised tokens per complete pass. Two epochs present about 40 million. At a measured 1,250 supervised tokens per second, that is 32,000 seconds, about 8.89 hours of processing. Add startup, evaluation, saving, preprocessing and interruptions. A rate that counts padding or prompt tokens is a different measure.

For a transformer dominated by square weight matrices, doubling width roughly quadruples many parameter terms. Doubling block count roughly doubles those terms. Embeddings, position tables, output heads, and architectural variations change the exact total. A good answer uses the actual configuration rather than a slogan about model size.

Splits and unavailable information

A strong split explanation names the independent unit and the intended generalization claim. Frames from one video, clips from one recording session, paraphrases of one source document, and photos of one physical part should usually remain grouped for the corresponding new-source evaluation. Counting derivatives as independent examples inflates apparent data diversity.

For a future-fault predictor, a measurement recorded after the fault cannot be an input to a prediction supposedly made before it. For an inventory assistant, today's stock level cannot be inferred reliably from last month's static training text alone. A larger model does not create missing current information. Appropriate responses include adding a legitimate input, retrieving current state, or abstaining.

A useful annotation policy resolves cases where reasonable people disagree. If a support message mentions both a late shipment and cancellation, specify the routing priority or allow multiple labels. Do not ask the optimizer to solve an undefined policy.

Classification and detection metrics

A detector finding 18 of 20 true targets with 12 false alarms has recall 18/20 = 0.90 and precision 18/30 = 0.60. Overall accuracy can look high when most images are empty. The relevant operating decision depends on missed-target cost, false-alarm workload, and whether abstention is available.

A box detector that returns a confident box in the wrong location has not made a correct detection. In the one-object exercise, a successful match requires both sufficient objectness and IoU at least 0.5. The objectness threshold and IoU criterion answer different questions.

For a mask that covers x coordinates 4 through 13 and y coordinates 2 through 5 in a 20 by 10 image, exclusive-edge coordinates are [4, 2, 14, 6]. Normalized xyxy is [0.2, 0.2, 0.7, 0.6]. This convention includes the full last foreground pixel. A mask with no foreground has no real object box, and its placeholder coordinates should not contribute to localization loss.

A large region shifted by one pixel can retain high IoU while its exact boundary differs substantially. Therefore an outline application should evaluate boundary tolerance in meaningful units, not only region overlap. An answer that proposes 'higher resolution' should also explain annotation precision, original-image coordinates, and compute cost.

Retrieval and contrastive learning

A useful hard negative is plausible but truly wrong for the query. A passage that paraphrases the correct answer is a false negative if treated as irrelevant. Review mined negatives before training; a similarity score alone does not establish irrelevance.

When a batch uses other positive passages as negatives, duplicate or semantically equivalent passages can contradict the training objective. Ordinary gradient accumulation over separate microbatches does not automatically create cross-microbatch negatives. A correct larger-batch proposal must explain how the loss sees the intended candidate set.

After retraining the encoder, regenerate document embeddings with the selected saved encoder and the same prompts, pooling and normalization. Reusing an old index with new query embeddings compares incompatible representations. The search program must load the matching encoder/index pair.

Language and typed decisions

A complete tool-call test checks whether an action is appropriate, whether the selected tool is correct, whether all arguments are valid and semantically correct, and whether the application permits execution. Parseable JSON solves only the first formatting layer. A model that fabricates a missing identifier should fail the task even if its schema is valid.

A typed-decision model returns scores over allowed choices. The output type limits the set of possible answers, not the chance of choosing the wrong one. Calibration must use separate examples, and an abstention policy must be judged by both accepted-case quality and coverage. A system that answers nothing is not automatically useful because its remaining error rate is low.

The causal-mask exercise changes a future input token and checks earlier-position logits in evaluation mode. Those earlier logits should remain unchanged within numerical tolerance. A low training loss cannot substitute for this information-flow test because an incorrect mask may make the task artificially easy.

A valid style evaluation preserves facts and compares observable style features on new topics. Copying familiar phrases from a small training collection is weaker evidence than adapting the style to genuinely new content. Blind comparisons reduce the temptation to reward the model you just trained.

Speech and generated audio

For reference 'turn the blue valve' and hypothesis 'turn blue valves', one plausible minimum edit alignment deletes 'the' and substitutes 'valves' for 'valve'. That is two errors over four reference words, so WER is 0.5. Normalization policy changes what counts as a word or substitution; record it and apply it consistently.

A transcript must match the exact audio segment. Cropping an utterance while retaining its complete transcript creates incorrect supervision. Keeping the same speaker or recording session across splits may also make an evaluation claim about unseen speakers invalid.

Speech recognition predicts text from recorded speech. A diffusion sound model generates audio from noise and conditions. Training a sound generator on tones does not teach speech transcription, and a transcription model does not automatically synthesize speech. A valid category choice begins with which direction the application needs.

For a denoising project, identify the precise target implemented by the trainer: noise, clean signal, velocity, or another parameterization. Preserve that relationship with the scheduler used at inference. A falling loss under one objective does not justify switching the sampler to a different unsupported convention.

The final independent project

A strong project plan contains the prediction contract, lawful data, independent split, baseline, training objective, small correctness test, measured feasibility test, fixed quality evaluation, artifact contract, and failure policy. It also states what evidence would make you stop.

If a first run fails, a strong next step identifies the earliest wrong boundary and proposes one controlled change. 'Use a bigger model' is justified only when evidence points to a capacity limit rather than missing information, bad labels, broken preprocessing, or an unsuitable objective.

When you can explain your model's input, target, learned values, loss, update, evaluation and deployment path without relying on the name of a library, you have completed the core learning goal of this book.