Training Your
Own Models
Browse the book
Chapter 38 / 4012 min read

38 Learn from public training projects without copying their assumptions

The most useful tutorial is one you can interrogate. After a small experiment works, open a second implementation and ask why its data, shapes, objective, evaluation, and saved files differ from yours. This makes public repositories and lessons useful throughout the book instead of turning them into a large prerequisite list. Use the book’s current-model recipes for practical work. Older lessons appear here only as small concept exercises or brief comparisons that explain why a current practice matters; their model recommendations and installation commands are not a current shopping list.

The reading path was checked on 2 October 2026 against repository code, configurations, tests, creator-written lesson material, and original YouTube descriptions. The video guidance relies on companion material and author-published chapter markers, not a claim of watching entire recordings or reading their caption transcripts. These sources were reviewed; their GPU workflows were not tested for this book.

Stage 1 Make a gradient visible

Start after the book's first small model and its explanation of one training step. You do not need a GPU to understand an update.

Read the small micrograd engine alongside its comparison tests. It represents scalar computations as a graph, orders that graph, and accumulates the contribution of every path into each gradient. The tests compare both outputs and gradients with PyTorch. This is a better first autograd exercise than implementing a large model whose wrong gradients can hide behind apparently plausible samples. T01, T02

Use Andrej Karpathy's micrograd lesson selectively: 08:08 for the first derivative, 1:22:28 for a reused-node bug, and 2:01:12 for the parameter-update walkthrough. The companion notebook explicitly corrects an exp backward assignment to accumulation. Prefer the corrected written implementation when it differs from a live recording. T03

Original exercise: evaluate y = x*x + x at x = 3. Its output is 12 and its derivative is 7. Check the derivative three ways: algebra, a small finite difference, and autograd. Next, deliberately overwrite rather than add a gradient contribution and explain why the answer changes. Then take a small step that reduces (y - 8)^2. This introduces derivative, gradient, graph, backward pass, and learning rate exactly when they become useful.

Grant Sanderson's gradient-descent lesson text supplies a visual bridge from one adjustable value to many. It distinguishes the prediction function from the cost function whose inputs are parameters. Use that distinction when a beginner asks whether training changes the input examples or the model. The text is an official adaptation credited to Josh Pullen; it is not a transcript read from the video. T04

Advance when: you can name which values are data, parameters, predictions, loss, and gradients; explain why gradients are reset between ordinary optimizer steps; and check one derivative independently. Scalar autograd is an educational microscope, not a GPU implementation strategy. “Small ML” here should not be confused with TinyML deployment on a microcontroller.

Stage 2 Turn text into supervised examples

Karpathy's makemore MLP notebook is valuable because it makes each next-character training pair visible. It shuffles whole names into training, development, and test partitions before producing overlapping context windows. Its embedding lookup, flattening, hidden layer, logits, and cross-entropy can be inspected separately. A shape shown in an exploratory notebook comment can become stale as a later cell changes dimensions: trust the actual tensor shape. T05

The associated MLP lecture is a follow-along companion, not something the beginner must finish before trying the project. The separate makemore executable repository packages multiple model families in one file and defaults to a tiny transformer. Do not assume that running the repository default reproduces the MLP lecture. T05, T06

Original exercise: use a small, permissioned list of invented product names. Split complete names first, build length-three context/next-character pairs, and print ten decoded pairs. Compare a count-based bigram against an MLP. Hold the split and evaluation code constant. Record held-out loss and samples, including malformed outputs and memorized names. This is the right moment to learn token, vocabulary, embedding lookup, context window, logit, cross-entropy, validation, and leakage.

Advance when: you can account for every row in a batch, explain what the target means, and explain why windows from the same source item must not casually cross dataset partitions.

Stage 3 Diagnose a network before enlarging it

In the activation-and-gradient notebook, the instructive artifacts are the activation histograms, saturation measurements, gradient histograms, and update-to-parameter ratios. Its running normalization statistics also show that a trained artifact can contain buffers as well as learnable parameters. These are reasons to inspect a model's state, not merely count its weights. T07

Original exercise: make a hidden layer's initial weights ten times larger, keep the data and seed fixed, and inspect the activation distribution and loss. Restore the initialization before testing a different learning rate. Record a hypothesis and one changed variable. Remove expensive debug hooks and retained intermediate gradients after the diagnostic run. Do not teach a numeric saturation percentage or update ratio as a universal pass/fail threshold.

Read normalization only after the experiment makes its purpose concrete. BatchNorm in this MLP lesson does not imply that a GPT should receive BatchNorm; the transformer architecture uses its own normalization design.

Stage 4 Replace fixed context with causal attention

The corrected GPT video companion shows shifted targets, a causal mask, head-sized query/key/value projections, multi-head concatenation, residual paths, and a pre-normalized block. Keep this as a read-only comparison unless its licensing is clarified: no license was reported by GitHub for this repository at inspection. Independently written book code is safer to distribute. T08

Karpathy's GPT lesson has particularly useful checkpoints at 14:27, 47:11, 1:02:00, and 1:26:48. They connect batching, weighted aggregation, learned attention, and residual connections. Its description corrects two mistakes: future tokens must not influence earlier predictions, and attention normalization uses the head dimension. These corrections were read directly on the author's YouTube page. T09

Original exercise: choose a short token sequence. Print the allowed attention positions as a grid; change only a future token; verify that earlier-position logits do not change in evaluation mode. Then compare a one-head implementation with a vectorized multi-head implementation using matching weights. A numeric equivalence test is stronger evidence than two runs both producing plausible text.

Sanderson's attention lesson text is helpful for visualizing a weighted update. It explicitly warns that its column-vector diagrams use a transposed convention relative to the original paper. Write tensor shapes beside every equation before porting it. Its adjective/noun example is an illustrative possible behavior, not a promise that every trained head has a clean human-readable purpose. T10

Now read sections 3.2.1-3.2.3 of Attention Is All You Need. Identify the scaling, learned projections, and masking in your own code. The original work is an encoder-decoder translation system; a small decoder-only GPT is not an exact reproduction. Save the original training recipe until you can interpret its eight-P100 hardware premise. T11

Advance when: you can trace [batch, time, channels], explain why targets are shifted once, distinguish self-attention from cross-attention, and pass the causal-leakage test.

Stage 5 Turn a notebook into a recoverable training run

For a current public process reference, inspect nanochat’s checkpoint manager together with the checkpoint block in base_train.py. Model weights, optimizer state, model configuration, run arguments, data-loader state, and loop state play different roles. The load path also distinguishes training from evaluation mode and checks tokenizer vocabulary size. This is a process-reading exercise; its default large run is not the book’s one-RTX command. T25

One historical comparison is useful: nanoGPT’s smaller checkpoint dictionary did not save every random/data-stream state. Its readable loop remains educational, but the repository is now explicitly deprecated. Do not use its old GPU timings or installation advice to plan a current run. T12, T13

Raschka's standalone training script offers another readable bridge. Its evaluation helper switches to evaluation mode and disables gradients before restoring training mode. It tracks tokens seen as well as optimizer steps. The short-story experiment is suitable for studying the loop and overfitting, not evidence that a small corpus creates a generally capable language model. T14

Original exercise: stop the same tiny run, reload it, and verify that evaluation logits agree before continuing. Separately test whether the resumed next optimization step matches an uninterrupted run. The first check verifies saved-model fidelity; the second requires optimizer, scheduler, random-state, and data-position accounting. Keep these claims separate in the experiment report.

Advance when: a fresh process can reconstruct the tokenizer and model, evaluate held-out data, generate with evaluation-mode settings, and explain the precise resume guarantee it does and does not provide.

Stage 6 Adapt a small pretrained model with auditable labels

Read Raschka's instruction collator before hiding label construction behind a trainer. It shifts targets and masks repeated padding while retaining a terminal token. Its default objective includes the instruction text: padding masking and assistant-only masking are different choices. That distinction is the lesson to transfer. T15

The Hugging Face smol-course LoRA lesson makes adapters, loading, and merging concrete. Pair it with the book's fixed small-model SFT recipe; do not import the course's entire changing environment into an already pinned setup. T16

Original exercise: before training, print one complete rendered conversation, its token IDs, attention mask, and loss mask. Color or label which tokens contribute to loss. Confirm that the answer has not been removed by truncation and that the chosen end-of-turn token remains supervised. Then overfit a tiny training-only slice as a plumbing test; a deliberately memorized slice is not the final evaluation.

The course's v2 hands-on page deserves a version audit. Its root requirements pin Transformers 4.46.3 and TRL 0.12.1, while the page uses newer SmolLM3-era examples, Trackio, and SFTConfig(max_length=...). The pinned TRL documentation instead uses max_seq_length. A later example passes args=config after defining training_config. Evaluation intervals alone do not create a held-out evaluation dataset. These are reasons to validate the precise chosen recipe, not evidence that the general method is wrong. T17, T18

Read LoRA once you can identify frozen weights, trainable weights, and the dimensions of a low-rank update. Its large-model result is evidence about the adaptation method, not proof of your 24 GB configuration's fit. The LLM chapter provides the actual memory and objective discussion.

Advance when: you can report trainable parameter count, supervised token count, adapter/base revisions, maximum sequence length, measured peak memory, a baseline comparison, and reload behavior. A tutorial's word “small” is not a memory budget.

Stage 7 Train an embedding for a real retrieval decision

The current Sentence Transformers NLI training example separates training loss from an embedding-similarity evaluator and evaluates the initial model. The MS MARCO example adds mined negatives, teacher-score filtering, a no-duplicates sampler, normalized vectors, and cached contrastive loss. The full example has substantial dataset and host-RAM requirements; its mini-batch control only addresses part of memory usage. T19

Original exercise: begin with your small retrieval corpus and the book's exact-search evaluator. Write ten difficult queries and label all known acceptable documents. Inspect the top errors before mining harder negatives. Verify that a supposedly negative document is not a second valid answer. Compare retrieval metrics before and after adaptation, then rebuild the corpus embeddings using the new encoder. Use the embedding chapter's pinned API rather than mixing its v5 recipe with a newer main-branch import layout.

This is where the meanings of embedding, pooling, bi-encoder, hard negative, false negative, in-batch negative, and Recall@k become operational. Read Sentence-BERT after observing why independently computed document vectors make retrieval possible. Do not require readers to retrain a generic embedding model from random weights before making a useful domain retriever.

Stage 8 Learn denoising before adapting an image or audio generator

Jonathan Whitaker's official diffusion lesson links the creator's video walkthrough and the actual Hugging Face notebooks. The scratch notebook deliberately starts with a tiny uniform-noise denoiser, then compares it with DDPM. Do not relabel the toy corruption rule as the DDPM forward process. Noise distribution, target parameterization, timestep conditioning, and sampler must be named separately. T20

Original exercise: on the book's synthetic small images, plot one clean example and several noised versions. Print the sampled timestep and target type. Verify the scheduler's forward-noise calculation numerically on one example. Only then compare learned denoising and generated samples on fixed seeds. Change one item at a time: model, target, schedule, or sampling steps.

The audio notebook is useful for tracing waveform → resampling → spectrogram → denoising → audio reconstruction. This particular example diffuses spectrogram images; it is not a general description of every audio diffusion architecture, and it does not teach transcription. Its loop contains loss.backward(loss). For ordinary scalar MSE training, use loss.backward(); the extra argument supplies an upstream gradient and multiplies the update signal by the loss value. T21, T22

Advance when: you can state what the model predicts, which representation it sees, how a sample is reconstructed, and which held-out checks measure quality beyond a loss plot. Start an audio round-trip test before training: if representation conversion already damages the source substantially, more denoising updates cannot simply erase that limitation. For advanced image/audio architectures and adaptation, follow the image and audio chapters and their primary papers.

Stage 9 Study faster implementations as controlled experiments

llm.c's LayerNorm study recomputes a normalized activation in backward and compares manual derivatives with autograd. Its repository also validates C results against a PyTorch reference. The transferable practice is “prove equivalence, then optimize,” not “rewrite everything in CUDA first.” Its CPU starter is a short pretrained GPT-2 fine-tune, and its legacy single-GPU FP32 path is distinct from its modern reproduction process. T23

nanochat extends the public process across tokenizer training, pretraining, SFT, evaluation, and inference. Its default speedrun targets an eight-H100 node. The README specifically warns that GPUs below 80 GB need configuration changes. Its CPU/MPS script is explicitly educational and shrinks the model. Do not convert either example's reported duration into an RTX estimate. T24

Original exercise: keep model outputs and evaluation fixed while comparing two attention implementations or precision settings. Report numerical tolerance, peak memory, tokens per second, warm-up policy, and evaluation difference. Decide whether the improvement is worth the added setup and debugging cost. Architecture research, pipeline engineering, and raw-kernel engineering are different next projects; the beginner need not complete all three to become a useful practitioner.

A compact source audit checklist

Before recommending a public command, inspect the actual file that will execute it.

  1. Record a commit or released version, not just a default-branch URL
  2. Read the data loader and collator before the optimizer settings
  3. Identify every download, login, telemetry integration, and upload
  4. Check whether “optional” publishing is actually executed: the inspected Sentence Transformers examples attempt push_to_hub inside try blocks
  5. Separate a source-code license from book artwork, weights, data, and generated-artifact rights
  6. Compare the tutorial's API names with its dependency file and the matching versioned docs
  7. Keep GPU model, GPU count, precision, sequence length, batch size, and token budget beside any runtime claim
  8. Run a tiny smoke test and a reload test before a longer job
  9. Read later corrections and tests as carefully as the tutorial narrative
  10. Mark every unrun GPU workflow honestly; a successful source review is not a benchmark