Training Your
Own Models
Browse the book
Chapter 31 / 405 min read

31 Diagnose failures before making the model larger

Locate the first wrong boundary

A training system has boundaries: raw input, decoded input, processed representation, batch, model output, loss, gradient, update, saved artifact, and deployed prediction. Find the earliest boundary where reality differs from expectation. This is usually more efficient than changing architecture in response to every poor result.

If the model's answer is wrong, first print the exact input it received. Was the relevant text truncated? Were image channels swapped? Was audio resampled correctly? Was a label ID mapped to the wrong name? Did an instruction template add a different generation prefix? The model may be responding correctly to an unintended input.

Keep a small diagnostic dataset whose examples you know by hand. It should include one ordinary case, one boundary case, one missing-input case, and one case requiring abstention. Run it after changes to preprocessing, dependency versions, export, or serving.

When loss does not fall

Try to overfit a tiny clean subset. If the model cannot reduce loss there, inspect whether the intended parameters require gradients, whether gradients are finite and nonzero, whether the optimizer receives those parameters, and whether optimizer.step is actually called. Check that zeroing gradients happens at the correct point.

Verify target alignment. A language model's logits at one position must be compared with the intended next token, not accidentally the same input token or a target shifted twice. A segmentation mask must match the transformed image. A classifier's label IDs must match the output dimension and class mapping.

Inspect the loss scale and supervised-token count. A mean over an empty mask is not a meaningful objective. A loss accidentally summed over very different lengths can produce unstable updates. A model output passed through softmax before a loss expecting raw logits can change the intended computation.

Only after those checks should you tune learning rate, initialization, or capacity. A larger network with the same alignment bug is a more expensive broken pipeline.

When training improves and validation worsens

This pattern suggests overfitting, but investigate its form. Are there duplicates in training? Is the dataset too small or unrepresentative? Are labels inconsistent? Has the model memorized an identifier? Did preprocessing differ between train and validation? Is the validation set from a meaningfully different distribution?

Possible responses include better data coverage, fewer updates, stronger justified regularization, a smaller trainable subset, or a better representation. The right response depends on the failure. If the deployment distribution changed, early stopping alone may not solve the mismatch.

Look at individual errors and slices. A model may improve on common short examples while regressing on long or rare cases. An average curve cannot tell you which data to add. An error taxonomy can.

When the GPU runs out of memory

Identify the phase. Failure during model loading concerns weights and initialization copies. Failure during forward often concerns activations, logits, or attention buffers. Failure during backward adds saved activations and gradients. Failure at the first optimizer step may reveal newly allocated optimizer states. Failure during evaluation can come from a different batch or generation cache. Failure during merge/export can come from temporary copies.

Reduce the relevant cost. Lower microbatch size or sequence length for activation pressure, enable a supported efficient attention path, use activation checkpointing, or select a smaller model. For persistent state, consider a supported optimizer policy, adapter training, or explicit offload. Do not delete arbitrary files or kill unrelated processes as a reflex; understand their ownership and purpose.

Check for retained computation graphs. Appending loss tensors to a Python list instead of detached scalar values can retain memory across steps. Keeping generated logits for every batch can do the same. A steadily rising memory curve may indicate a leak rather than a model that is intrinsically too large.

When the GPU is mostly idle

The bottleneck may be data loading, tokenization, image decoding, augmentation, disk access, CPU preprocessing, or synchronization. Measure time by stage. Preprocess deterministic work once when appropriate. Use a reasonable number of data-loader workers and avoid overwhelming RAM or storage. Cache only what remains valid under the chosen augmentation and model settings.

A tiny model can also be launch-bound: the GPU spends a large fraction of time starting small operations. Larger batches may improve utilization if memory allows. Compilation can help some workloads but adds startup cost and compatibility complexity. Establish a correct baseline before optimizing it.

Do not compare throughput numbers without their units and shapes. Images per second at 256 pixels and at 1024 pixels describe different work. Tokens per second can count padded tokens, input tokens, supervised tokens, or generated tokens. Name the measure.

When outputs look memorized or repetitive

Check data duplication, excessive epochs, narrow prompt coverage, generation settings, and stopping tokens. A style adapter trained on a tiny collection can reproduce phrases rather than learn general style. An image adapter can bind a subject to its background. Audio training can overfit a repeated segment or generate a narrow set of textures.

Evaluate on prompts designed to separate the intended concept from accidental correlations. Ask for the subject in a different scene. Ask for the same writing style on a new topic. Ask for a different sound duration or context within the model's supported interface. If the model fails, add representative examples or reduce the learned change; do not simply select a more flattering seed.

When the exported model differs

Compare the training checkpoint, unmerged adapter, merged full-precision model, and quantized export on the same fixed inputs. Confirm identical tokenizer, chat template, special tokens, context settings, preprocessing, and generation parameters. Check that evaluation mode is active.

Quantization can change outputs and may have format-specific operator support. A successful conversion is not a quality test. An export that loads in a CPU runtime may still use a different tokenizer or fail on long inputs. Validate the actual deployment path.

A debugging notebook that compounds in value

For each failure, record the symptom, smallest reproducer, expected behavior, observed behavior, diagnosis, change made, and verification. Save the exact error message without secrets. Distinguish a confirmed cause from a plausible hypothesis.

A good failure note prevents rediscovering the same problem. It also protects you from attributing every improvement to the most interesting change. Sometimes the decisive fix was correcting a label mapping rather than adopting a new architecture.

Exercises

Introduce an intentional target shift in a copy of the tiny transformer and observe the effect. Restore the original before continuing. Explain why a pipeline can run successfully while optimizing the wrong task.

Create a fresh-process inference test for one saved model and compare it with in-process output. Then change one preprocessing setting and document the mismatch.

Write an out-of-memory decision tree organized by the phase of failure. For each branch, name one measurement you would gather before changing a setting.