Training Your
Own Models
Browse the book
Chapter 14 / 404 min read

14 Make a run reproducible and recoverable

Save the experiment rather than only the weights

A weight file is not a complete experiment. You also need architecture, tokenizer or feature mapping, preprocessing, data version, train/validation/test assignments, optimizer configuration, training duration, and inference settings. Without these, a future version of you may have a model that cannot be loaded correctly or evaluated fairly.

For each run, save a resolved configuration rather than only command-line overrides. If a default changes in the library, the resolved value should remain visible. Record Python and package versions, GPU model, driver, numeric policy, attention backend when known, code revision, model revision, and hashes of input splits.

A model repository name can point to changing contents. Resolve and retain a commit revision. A GitHub branch is similarly mutable. A tagged release is better for reproducibility, and an immutable commit hash is stronger evidence of exactly what was used. For a published training process, archive the relevant configuration and record local modifications.

Randomness has several sources

Initialization, shuffling, dropout, augmentation, sampling, and some GPU operations involve randomness or nondeterministic execution. Set the seeds used by Python, NumPy when present, PyTorch, and data-loader workers as appropriate. A dedicated random generator for dataset sampling can make its state easier to preserve.

A seed does not guarantee identical results across software releases, devices, or platforms. Some algorithms are nondeterministic, and floating-point order can change. PyTorch explicitly cautions that complete cross-platform reproducibility is not guaranteed and that deterministic alternatives can cost performance. F14

For a serious comparison, repeat the selected small experiment with more than one seed when affordable. If a claimed improvement disappears under a different initialization or shuffle, report its instability. For a first correctness exercise, deterministic fixed seeds are useful because they reduce distractions.

Checkpoints serve different purposes

An inference checkpoint contains enough to make predictions: weights, architecture configuration, tokenizer or preprocessing, and any required adapter or head metadata. A training-resume checkpoint additionally needs optimizer state, scheduler state, update number, gradient-scaler state where used, and relevant random and sampler states.

An adapter checkpoint ordinarily does not contain the base model. Record the exact base identity and revision. Loading an adapter on a similarly named but different checkpoint may fail or silently change behavior. A merged export is a separate artifact with its own validation requirements.

Checkpoint interval is a tradeoff. Saving too often consumes time and disk; saving too rarely risks losing work. Estimate the amount of recomputation you are willing to tolerate and choose an interval accordingly. Keep a last recoverable checkpoint and selected evaluation checkpoints. A retention policy should not delete the only known-good artifact before a replacement has been checked.

Practice interruption on a tiny run

Start a short run that saves twice. Stop it after the first save, resume, and verify that the next recorded step continues appropriately. Compare configuration and data hashes. Confirm that an intentionally mismatched architecture or dataset is rejected.

A resumed run may differ from an uninterrupted run if data-loader state, random generators, or batch order were not restored. Be honest about that limit. 'Resumes without crashing' and 'continues the identical optimization trajectory' are different claims.

Use atomic or transactional saving when possible: write to a temporary location, complete the write, then replace the designated checkpoint. Check free disk space. A full filesystem can produce a truncated artifact after hours of otherwise successful training.

Logs should answer operational questions

At minimum, log update number, training loss, learning rate, evaluated validation metric, elapsed time, and throughput with clear units. For GPU work, include peak memory. Useful additional signals include gradient norm, supervised-token count, skipped nonfinite updates, truncation count, and examples per class or source.

Do not write every private training example into a public experiment tracker. Some frameworks enable external reporting by default. Read the configuration for report_to, push_to_hub, telemetry, and upload behavior. Disable unneeded external reporting for a local experiment, and obtain appropriate permission before transmitting private data or model artifacts.

Store a small, sanitized set of fixed qualitative evaluation outputs alongside metrics. Numbers can show that something changed; outputs often explain what changed. Keep the exact prompt, seed, decoding settings, and reference answer with each sample.

A minimal run ledger

Use a record with fields for run name, question being tested, baseline, data version, model revision, code revision, environment file, configuration path, start and end time, status, measured memory, measured throughput, primary metric, guardrails, selected checkpoint, and next decision.

A failed run still receives a record. State whether it failed before loading, during the first forward pass, during backward, at the optimizer step, during evaluation, or during export. That stage narrows the diagnosis. Do not erase failures and then claim the workflow was effortless or fully tested.

Exercises

Take one checkpoint and load it in a fresh process with no variables left from the trainer. If you cannot reconstruct preprocessing and output decoding from saved files, add the missing metadata.

Change a dataset file by one character and confirm its hash changes. Then explain why a filename is not a reliable dataset version.

Review a public training script for external reporting, automatic uploads, downloads, remote-code execution, overwrite behavior, and unpinned dependencies before running it. A published example is evidence to inspect, not authority to execute everything it contains.