Training Your
Own Models

Change one thing. Inspect the result.

Three ideas, made tangible.

Small browser-only illustrations of concepts in the book. These are educational calculations, not GPU training runs or hardware benchmarks.

01 / Learning from a loss

Watch a parameter learn.

Our entire toy model is one number, w. Its target is 3, and its loss is (w − 3)². Each step uses gradient 2(w − 3), then updates w ← w − learning rate × gradient.

Read chapter 3 →

The chart rescales to the largest observed loss. The numbers above show absolute values. Reset starts at w = −2. The controls stop at 150 steps; reset to begin again. This convex one-parameter example does not represent a neural network’s full loss landscape.

02 / State before activations

Where does the memory go?

This calculation inventories persistent full-training state for a dense model. It excludes activations, attention workspaces, temporary copies, allocator overhead, frozen components, and display use. It cannot determine whether a real run fits.

Read chapter 15 →

FP32 case: 4 + 4 + 8 = 16 bytes per parameter. Mixed-storage case: 2 + 2 + 4 + 8 = 16 bytes per parameter. These are deliberately explicit assumptions, not universal framework defaults. GiB = 2³⁰ bytes; decimal GB = 10⁹ bytes.

03 / Information is directional

Which positions can this token see?

Select a query position in this six-token sequence. In an inclusive causal mask it can attend to itself and earlier positions; future positions are blocked.

Read chapter 12 →

This shows allowed positions, not learned attention weights. Causal masking governs access to context. A loss mask separately selects which targets contribute to the training objective.

Changes stay in this page and reset when it reloads. No data is uploaded.