Training Your
Own Models
Browse the book
Chapter 6 / 406 min read

6 Decide whether the model actually improved

Make the baseline answer first

Before training, run the simplest plausible system on the evaluation cases. For a classifier, this might be the most frequent class, a short rule list, or a TF-IDF model. For a pretrained LLM, it is the untouched checkpoint with a careful prompt. For retrieval, it may be keyword search or the unmodified embedding model. For diffusion, use the base model with fixed prompts and seeds.

Save the baseline outputs. Otherwise a newly trained model's fluency or visual novelty can make it seem better than a baseline you remember inaccurately. Evaluation is a comparison under a fixed procedure, not a collection of favorite examples.

Write a primary metric and a small number of guardrails. The primary metric expresses the improvement you are trying to buy. Guardrails catch unacceptable regressions, such as malformed JSON, excessive latency, privacy leakage, poor performance on an important minority slice, or catastrophic forgetting of a required capability.

Classification needs more than one accuracy number

A confusion matrix counts how often each true class was assigned each predicted class. It shows which categories are confused. Precision asks, among cases predicted positive, how many were actually positive. Recall asks, among actual positives, how many were found. F1 combines precision and recall into a single harmonic mean, but its convenience does not remove the need to inspect the underlying tradeoff.

Suppose a defect detector finds 18 of 20 true defects and raises 12 false alarms. Recall is 18/20 = 0.90. Precision is 18/30 = 0.60. Whether that is acceptable depends on the cost of missed defects and false alarms. A high overall accuracy on thousands of normal cases could conceal the operational problem.

For multiple classes, macro averaging gives each class equal weight; micro averaging combines decisions across all examples. Weighted averaging uses class prevalence. State which you use. An apparently improved average may hide deterioration on a rare but important class.

For numeric regression, inspect mean absolute error, root mean squared error where relevant, and errors by target range. A model can have a good average error while failing badly on unusually large values. Compare against a constant mean or median baseline, and make sure the metric's units are understandable.

Generated outputs need decomposed tests

For a structured answer, first test whether it parses. Then test schema validity, required fields, allowed values, and semantic correctness. Valid JSON containing the wrong account ID is still wrong. A tool call should be evaluated for tool selection, argument extraction, missing-information handling, and whether an action should have been proposed at all.

For a knowledge answer, separate evidence retrieval from answer generation. Did the relevant passage enter the context? Did the answer accurately use it? Did it cite the right source? Did it invent unsupported details? A failed answer can originate in retrieval, context construction, the model, or the output validator.

For writing style, use a rubric that describes observable features rather than 'sounds like me'. Examples include typical sentence length, degree of formality, use of headings, tolerance for hedging, and factual preservation. Ask evaluators to compare outputs blind to which system produced them. Keep content quality separate from style similarity.

For images, compare identity or subject fidelity, prompt adherence, diversity, unwanted copying, and artifacts across fixed prompts and seeds. For audio, listen for event identity, timing, clipping, noise, repetition, and intelligibility where speech is involved. A numerical proxy is useful when it correlates with the task, but it should not replace listening or viewing.

Test behavior outside the easiest distribution

Include ordinary cases, edge cases, missing inputs, contradictory inputs, irrelevant content, and cases requiring refusal or abstention. Test changes in wording that should preserve the answer. Test changes in facts that should change it. These paired tests can reveal whether the model learned the intended dependency.

For a tool model, change a quantity or date in the input and see whether the corresponding argument changes while unrelated fields remain stable. For a classifier, remove an accidental keyword and test whether the decision still follows the substance. For image generation, change the requested background and see whether the learned subject is entangled with its training backdrop.

Do not use adversarial cases only as a dramatic final demonstration. Design a bounded suite before training. If an error matters in deployment, it deserves a place in the evaluation contract.

Small samples have uncertainty

If a model succeeds on 18 of 20 cases, the observed rate is 90 percent, but the underlying success rate is not known precisely. One extra error changes the score by five percentage points. Reporting many decimal places does not create certainty. Record counts, sample selection, and important slices alongside percentages.

When comparing two models, evaluate them on the same cases. A paired comparison can reveal which cases improved and which regressed. For human ratings, randomize presentation order and use consistent instructions. If possible, use more than one evaluator on a subset and investigate disagreements.

Repeatedly choosing the best result on one validation set can overfit the evaluation procedure. Keep the final test set untouched while choosing hyperparameters. After a final test reveals a problem, it is legitimate to improve the system, but that test has then informed development. A new independent test is needed for a fresh unbiased claim.

Calibration and abstention

A predicted probability of 0.9 does not automatically mean the model is right 90 percent of the time on similar predictions. Calibration measures how probabilities correspond to observed outcomes. A classifier can rank examples well while being overconfident. Generative token probabilities are especially easy to mistake for answer-level confidence.

Choose abstention rules using held-out data and the actual cost of errors. For example, route low-confidence support tickets to a person. Measure both coverage, the fraction answered automatically, and accuracy among answered cases. A system can improve apparent accuracy by refusing almost everything; that tradeoff must be visible.

Temperature scaling for a classifier adjusts logits using a scalar fitted on a calibration set. This is different from freely raising generation temperature for creative text. Calibration requires separate data and should be evaluated on still-held-out cases. The typed-decision chapter develops this more carefully.

Loss and task quality can disagree

Training loss is useful for debugging optimization. Validation loss is useful for detecting some forms of overfitting or distribution mismatch. Neither replaces task evaluation. A language model can reduce average token loss by learning common formatting while still choosing the wrong tool. A diffusion model can improve its denoising objective while producing less diverse samples.

An instructive experiment keeps both curves and task outputs. At selected checkpoints, run the fixed evaluation suite and preserve outputs. Choose the checkpoint by the intended task and guardrails, not automatically by the last training step. If validation loss improves while task quality worsens, investigate label format, weighting, generation settings, and whether the loss actually matches the goal.

An evaluation report you can use

A concise report names the task, dataset version, split rule, baseline, candidate, primary metric, guardrails, generation or inference settings, measured latency and memory where applicable, counts, and failure examples. It states which cases remain unsupported. It links the exact model artifact and configuration.

Do not hide bad examples by averaging them away. A short error taxonomy is often the most useful output: wrong label, missing field, unsupported fact, truncation, failure to abstain, memorized background, or poor performance on long inputs. The taxonomy tells you what to change next.

Exercises

Write a metric that could look good while your intended system fails. Then add a guardrail that would reveal the failure. For example, overall routing accuracy can hide poor recall on urgent cases; a slice-specific recall requirement makes that visible.

Create five paired tests where a small input change should change exactly one part of the output. Create five more where harmless rewording should not change the decision. Run both sets on the baseline before training.

For your project, define the cost of a false positive, a false negative, and an abstention. If these costs cannot be expressed in money, describe their practical consequences. Use that description to choose a meaningful operating threshold.