Change one thing. Inspect the result.
Three ideas, made tangible.
Small browser-only illustrations of concepts in the book. These are educational calculations, not GPU training runs or hardware benchmarks.
01 / Learning from a loss
Watch a parameter learn.
Our entire toy model is one number, w. Its target is 3, and its loss is (w − 3)². Each step uses gradient 2(w − 3), then updates w ← w − learning rate × gradient.
The chart rescales to the largest observed loss. The numbers above show absolute values. Reset starts at w = −2. The controls stop at 150 steps; reset to begin again. This convex one-parameter example does not represent a neural network’s full loss landscape.
02 / State before activations
Where does the memory go?
This calculation inventories persistent full-training state for a dense model. It excludes activations, attention workspaces, temporary copies, allocator overhead, frozen components, and display use. It cannot determine whether a real run fits.
FP32 case: 4 + 4 + 8 = 16 bytes per parameter. Mixed-storage case: 2 + 2 + 4 + 8 = 16 bytes per parameter. These are deliberately explicit assumptions, not universal framework defaults. GiB = 2³⁰ bytes; decimal GB = 10⁹ bytes.
03 / Information is directional
Which positions can this token see?
Select a query position in this six-token sequence. In an inclusive causal mask it can attend to itself and earlier positions; future positions are blocked.
This shows allowed positions, not learned attention weights. Causal masking governs access to context. A loss mask separately selects which targets contribute to the training objective.
Changes stay in this page and reset when it reloads. No data is uploaded.