Training Your
Own Models
Browse the book
Reference37 min read

Works cited

Checked 2 October 2026 unless a source entry states otherwise. A cited implementation can explain a design without being a measured reproduction. Repository main branches can change; use the pinned versions and captured revisions in each project.

Foundation

  • F08. Hoffmann, J. and colleagues. Training Compute-Optimal Large Language Models. 2022. Empirical scaling of parameters and data under compute budgets; not a universal prescription for every domain. https://arxiv.org/abs/2203.15556

Tiny Ml Embedding

  • M02. scikit-learn 1.8, Probability calibration. Supports reliability interpretation, independent calibration data, FrozenEstimator, sigmoid calibration, and limitations of Brier/log-loss interpretation
  • M05. PyTorch 2.8, BCEWithLogitsLoss and CrossEntropyLoss. Logit/target conventions and stable loss APIs; the teaching Dice term is implemented explicitly in the accompanying script
  • M06. torchvision 0.23, MobileNetV3-Small. Explicit weight enum and associated transforms
  • M09. OpenCV, Structural Analysis and Shape Descriptors. findContours, hierarchy/retrieval modes, and polygon-approximation semantics. This older versioned reference documents the classic API used by the optional contour recipe
  • M10. PyTorch, TorchVision Object Detection Finetuning Tutorial. Dataset contract, box/instance masks, label conventions, and pretrained Mask R-CNN fine-tuning. The tutorial states torchvision >=0.16 for its current variant; follow the pinned 0.23 API when combining with this book
  • M15. Sentence Transformers, Samplers. Duplicate-avoidance rationale; the book's direct loop performs its own one-document-per-batch grouping
  • M18. Sentence Transformers, Evaluation reference. Information-retrieval evaluator terminology and supported metrics; this book includes a small independent exact-search metric implementation
  • M20. scikit-learn, Cross-validation guide. Group-aware and time-aware evaluation principles; this source was the stable guide at verification, while the executed estimator version is 1.8.0
  • M21. torchvision, official compatibility table. PyTorch 2.8 pairs with torchvision 0.23 and supports Python 3.9-3.13 according to the published table
  • M24. scikit-learn, Model persistence. Security and compatibility cautions for pickle-derived model artifacts
  • M26. Sentence Transformers, v5.1.1 release. The named version exists; this is an API target, not a tested environment lock
  • M28. torchvision v0.23.0, official detection training process and reference scripts. Published scripts, losses/evaluation support, and documented original training commands. Many original commands use eight GPUs; they are process references, not recipes promised to fit one 24 GB GPU unchanged
  • M29. Sentence Transformers, official embedding training examples. Public examples for retrieval, similarity, triplet, and other supervision styles; main-branch code can differ from the pinned v5.1.1 direct-loop recipe

Current Embedding

  • M30. Qwen, Qwen3-Embedding-0.6B official model card. Public/non-gated at verification, Apache-2.0 metadata, 0.6B family label, 1024-dimensional output, 32K stated context, minimum library requirements, and recommended query-only instructions. No leaderboard claim is used as an application-quality guarantee

Llm Training

  • L02. Qwen3-0.6B official model card. URL: https://huggingface.co/Qwen/Qwen3-0.6B Inspected: Model Overview; Quickstart; thinking/non-thinking mode; Apache-2.0. Card says 0.6B, 28 layers, 16 Q/8 KV heads and 32,768 context. No quality benchmark copied
  • L03. Qwen3-0.6B official Hub metadata. URL: https://huggingface.co/api/models/Qwen/Qwen3-0.6B Inspected: Inspected sha c1899de289a04d12100db370d81485cdf75e47ca, tokenizer template, 751,632,384 serialized BF16 entries. This is not a unique trainable-parameter count
  • L05. TRL supervised fine tuning trainer. URL: https://huggingface.co/docs/trl/v1.14.1/sft_trainer Inspected: v1.14.1; expected dataset formats; completion/assistant masks; training templates; tool columns; default chunked_nll. Used for API review, not a copied runnable benchmark
  • L14. Hu et al LoRA. URL: https://arxiv.org/html/2106.09685v2 Inspected: arXiv v2, 2021; low-rank reparameterization and frozen pretrained weights. Matrix example in book calculated independently
  • L18. Dettmers et al QLoRA. URL: https://arxiv.org/html/2305.14314v1 Inspected: arXiv v1, 2023; frozen four-bit base, NF4, double quantization and paged optimizers. No paper hardware result represented as a local measurement
  • L21. Lewis et al Retrieval Augmented Generation. URL: https://arxiv.org/abs/2005.11401 Inspected: Abstract and publication metadata; parametric generation combined with non-parametric retrieval memory
  • L22. Gururangan et al Do not Stop Pretraining. URL: https://arxiv.org/abs/2004.10964 Inspected: Abstract and publication metadata; domain/task adaptation in studied settings. Not treated as a universal modern-decoder guarantee
  • L23. Meng et al Locating and Editing Factual Associations in GPT. URL: https://arxiv.org/abs/2202.05262 Inspected: Abstract and metadata; ROME factual-association editing scope. No unsupported empirical locality rate used
  • L24. Meng et al Mass Editing Memory in a Transformer. URL: https://arxiv.org/abs/2210.07229 Inspected: Abstract and metadata; MEMIT many-edit research scope. Engineering cautions in book are recommendations, not claimed study results
  • L27. Berkeley Function Calling Leaderboard. URL: https://gorilla.cs.berkeley.edu/leaderboard.html Inspected: Inspected V4 methodology links; snapshot says models evaluated at f7cf735 and bfcl-eval==2025.12.17. No leaderboard number reported as an author measurement
  • L29. Rafailov et al Direct Preference Optimization. URL: https://arxiv.org/abs/2305.18290 Inspected: arXiv v3 metadata and abstract; reference-relative pairwise objective checked against TRL math in L30
  • L31. Hinton et al Distilling the Knowledge in a Neural Network. URL: https://arxiv.org/abs/1503.02531 Inspected: Abstract and metadata; teacher distributions and distillation origin. Temperature explanation is standard derivation, not a reproduced result
  • L32. Kim and Rush Sequence Level Knowledge Distillation. URL: https://arxiv.org/abs/1606.07947 Inspected: Abstract and metadata; sequence-level teacher targets distinguished from logit matching
  • L35. Qwen3-1.7B official model card. URL: https://huggingface.co/Qwen/Qwen3-1.7B Inspected: Model Overview and license; nominal 1.7B, 28 layers, 16 Q/8 KV GQA heads and 32,768 published context
  • L36. Qwen3.5-0.8B official model card. URL: https://huggingface.co/Qwen/Qwen3.5-0.8B Inspected: Model Overview; LM parameters separate from vision encoder, 24 layers, DeltaNet/gated attention layout, padded vocabulary 248320; Apache-2.0
  • L37. Qwen3.5-2B official model card. URL: https://huggingface.co/Qwen/Qwen3.5-2B Inspected: Model Overview; 2B LM label, vision encoder, 24 layers, hybrid layout. No benchmark transferred to local hardware
  • L39. Gemma 3 1B instruction model card. URL: https://huggingface.co/google/gemma-3-1b-it Inspected: Model Information, access conditions, context, license label. Family-level multimodal text not applied indiscriminately to 1B
  • L49. SmolLM3 3B official model card. URL: https://huggingface.co/HuggingFaceTB/SmolLM3-3B Inspected: Training/software/hardware sections; 11T pretraining tokens and 384 H100. Used as published process scale, not practical full-training project
  • L51. Gemma 4 E2B instruction model card. URL: https://huggingface.co/google/gemma-4-E2B-it Inspected: Dense-model table and license; 2.3B effective, 5.1B with embeddings, per-layer embeddings, Apache-2.0. Excluded from strict total scope
  • L54. Hoffmann et al Training Compute Optimal Large Language Models. URL: https://arxiv.org/html/2203.15556v1 Inspected: arXiv v1; section 3 and efficient-frontier 6ND approximation, tables/context. Book 1B/20B-token/50TFLOP example is explicitly hypothetical
  • L56. LFM2.5 350M official model card. URL: https://huggingface.co/LiquidAI/LFM2.5-350M Inspected: Model Details, Tool Use, Fine-Tuning, exported formats; 16 blocks, 32K context, vocab=65536, custom LFM1.0 label. Vendor speed claims not reused as local measurements
  • L57. Liquid official TRL fine-tuning guide. URL: https://docs.liquid.ai/lfm/fine-tuning/trl Inspected: LoRA Fine-Tuning and full-update examples; attention projections explicitly targeted; inspected old tokenizer keyword and minimum-version install string. Current API correction comes from L05
  • L59. Liquid official llama.cpp deployment guide. URL: https://docs.liquid.ai/deployment/on-device/llama-cpp Inspected: GGUF downloading, CPU-first execution, CLI/server and GPU-offload distinctions; model-specific existing export route, not proof of arbitrary adapter conversion
  • L60. LFM2 technical report. URL: https://arxiv.org/abs/2511.23404 Inspected: arXiv v1 abstract/metadata; hybrid gated-short-convolution/GQA design and hardware-in-the-loop search. No claimed reproduction of reported CPU gains

Jev Architecture

Diffusion Training

  • D01. Ho, Jain and Abbeel, Denoising Diffusion Probabilistic Models (2020). https://arxiv.org/abs/2006.11239 - foundational Gaussian forward corruption, noise-prediction training and reverse generation. Used for the short objective explanation; our waveform code is an original teaching implementation, not a reproduction of their benchmark results.
  • D02. Song et al., Score-Based Generative Modeling through Stochastic Differential Equations (2020/2021). https://arxiv.org/abs/2011.13456 - score interpretation and reverse SDE/probability-flow viewpoint. No paper benchmark is presented as a result of our scripts.
  • D03. Lipman et al., Flow Matching for Generative Modeling (2022/2023). https://arxiv.org/abs/2210.02747 - vector-field regression along probability paths; distinguishes flow matching from a particular epsilon-prediction recipe.
  • D04. Ho and Salimans, Classifier-Free Diffusion Guidance (2022). https://arxiv.org/abs/2207.12598 - conditional/unconditional predictions and inference-time quality/diversity trade-off. The chapter clearly states its guidance-scale convention.
  • D05. Rombach et al., High-Resolution Image Synthesis with Latent Diffusion Models (2021/2022). https://arxiv.org/abs/2112.10752 - separates an autoencoder representation from conditional latent diffusion.
  • D06. Peebles and Xie, Scalable Diffusion Models with Transformers (2022/2023). https://arxiv.org/abs/2212.09748 - transformer backbone on latent patches; architecture is distinct from output target.
  • D07. Ruiz et al., DreamBooth (2022). https://arxiv.org/abs/2208.12242 - subject-driven personalization and class-specific prior preservation; not an alternative mathematical definition of LoRA.
  • D17. Nichol and Dhariwal, Improved Denoising Diffusion Probabilistic Models (2021). https://arxiv.org/abs/2102.09672 - source for the cosine cumulative-noise schedule idea. The supplied code fixes s=0.008 and clips beta at 0.999; it does not implement every improvement in the paper.
  • D18. Kong et al., DiffWave: A Versatile Diffusion Model for Audio Synthesis (2020/2021). https://arxiv.org/abs/2009.09761 - waveform diffusion and conditional/unconditional audio synthesis. Our U-Net architecture is explicitly not described as DiffWave.
  • D16. Hugging Face Datasets v3.6.0, Create an image dataset. https://huggingface.co/docs/datasets/v3.6.0/en/image_dataset - imagefolder and JSONL metadata relationship. Our extra group field and split discipline are deliberate project design, not a feature that automatically prevents leakage in the upstream trainer.
  • D24. Stable Audio 3 actual training script. https://github.com/Stability-AI/stable-audio-3/blob/main/scripts/train_lora.py - inspected parser and training code, including base-model-only loading; raw audio plus .txt metadata; duration, rank, batch, local CSV logging; and demo sample-size handling. --save_dir is the actual parser name at the inspected source; do not blindly copy prose mentioning --output_dir. Its full dependency stack was not installed or executed for the book.
  • D26. Stable Audio 3 Small SFX official model card. https://huggingface.co/stabilityai/stable-audio-3-small-sfx - describes the family, base relationship, T5Gemma conditioning and additional Gemma terms. Note that the card's lower-level sample has a CPU-path variable issue (model_half is not initialized in the CPU branch); the book does not copy that snippet or claim parity with the higher-level API.
  • D28. Original DDPM implementation. https://github.com/hojonathanho/diffusion - primary author repository. Its README specifies TensorFlow 1.15, Python 3.5 and TPU experiments. Read to connect the paper with its historical implementation; do not present its installation as a modern single-RTX recipe.
  • D29. DiffWave reference implementation. https://github.com/lmnt-com/diffwave - public waveform/vocoder training and inference implementation. Useful next reading after the toy class-conditioned example; a mel-conditioned vocoder's data contract is different from text-conditioned sound generation.
  • D32. Actual Diffusers v0.36.0 Sana inference pipeline. https://github.com/huggingface/diffusers/blob/v0.36.0/src/diffusers/pipelines/sana/pipeline_sana.py - verified SanaPipeline, LoRA loading mixin, CPU-offload sequence, prompt embeddings and masks, 32-divisible dimensions, sampling inputs and decoder postprocessing. Training and inference explicitly agree on max_sequence_length=128 and complex_human_instruction=None instead of inheriting the pipeline's different instruction-prefix default.
  • D33. Chen et al., SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers (2024/ICLR 2025). https://arxiv.org/abs/2410.10629 - efficient architecture and representation design. Its reported inference hardware/result is not represented as a measured fine-tuning result for the book.
  • D35. Official maintained NVlabs Sana repository. https://github.com/NVlabs/Sana - current family, training/inference documentation and model variants. It is actively maintained beyond the original paper; the book selects a stable checkpoint with an explicit fine-tuning route rather than claiming the newest family release automatically has the same trainer.

Cohere Asr

  • S01. Qwen3-ASR-0.6B official card. URL: https://huggingface.co/Qwen/Qwen3-ASR-0.6B Inspected: 2026 release identity, Apache-2.0, language identification and 30-language/22-dialect coverage, separate forced aligner, inference wrapper, submodel versus whole-model size distinction. No vendor throughput result is adopted as a consumer-GPU measurement.
  • S03. Qwen3-ASR technical report. URL: https://arxiv.org/abs/2601.21337 Inspected: official report abstract and model-family scope; no replication of the training corpus or benchmarks is claimed.
  • S04. Cohere Transcribe 03-2026 official card. URL: https://huggingface.co/CohereLabs/cohere-transcribe-03-2026 Inspected: 2B, Apache-2.0, 14 languages, observed contact-sharing gate, native Transformers >=5.4.0 route, reported PyTorch 2.10.0 testing, missing timestamps/diarization, language and non-speech limitations. No access gate accepted. Weights not downloaded.
  • S05. Cohere technical announcement. URL: https://cohere.com/blog/transcribe Inspected: Conformer encoder + Transformer decoder, waveform/log-Mel input, supervised token cross-entropy and from-scratch original training. These are model characteristics, not evidence that full training fits 24 GB.
  • S08. North Micro Vision Instruct official card. URL: https://huggingface.co/CohereLabs/North-Micro-Vision-Instruct Inspected: 2.4B total = 2B language + 400M vision; Apache-2.0; model-specific Transformers 5.16.0 support and linked training recipes. Used only to place the model in the right category and avoid treating a 2B language-backbone label as a total count.
  • S09. Tiny Aya official card. URL: https://huggingface.co/CohereLabs/tiny-aya-global Inspected: explicit model-summary count 3.35B and CC-BY-NC-4.0. The card contains inconsistent boilerplate elsewhere; we rely on the model summary, not the erroneous 111B sentence in its terms section. No license permission beyond the card is inferred.
  • S11. Conformer foundational paper. URL: https://arxiv.org/abs/2005.08100 Inspected: convolution-augmented attention architecture concept. Historical reference for the mechanism, not an outdated checkpoint recipe.
  • S13. Official Cohere fine-tuning repository. URL: https://github.com/cohere-ai/cohere-finetune Inspected: supported base-model list and text-model LoRA/QLoRA scope. It does not establish Transcribe support. Mutable main inspected, not installed.
  • S17. CTC paper. URL: https://www.cs.toronto.edu/~graves/icml_2006.pdf Inspected: original connectionist temporal classification construction, blank labels and sum over valid alignments. Historical objective contrast only; the chapter does not claim its Qwen or Cohere loop uses CTC.
  • S22. PyTorch 2.10 automatic mixed precision. URL: https://docs.pytorch.org/docs/2.10/amp.html Inspected: autocast context and dtype behavior. The original loop explicitly uses FP32 model/optimizer storage and BF16 CUDA autocast. No GradScaler is used for that BF16 path; no claim of FP16 compatibility is made.
  • S24. Official PyTorch wheel installation matrix. URL: https://pytorch.org/get-started/previous-versions/ Inspected: PyTorch 2.10.0 Linux/Windows CUDA 12.6/12.8/13.0 and CPU wheel indexes. The chapter selects only torch for this project; torchvision/torchaudio are not required by its code. No package installation executed.

Large Architecture

  • A01. DeepSeek-AI. Official DeepSeek-V3 inference implementation, especially MLA, Gate, Expert and MoE. Inspected source; file revision b15f0dbbbe6a4bc403306175698439ef380f5fb5, dated 27 August 2025. The optimized absorb path stores kv_cache and pe_cache; the code also exposes a naive path. Pinned implementation
  • A02. XiaomiMiMo. MiMo-V2.6-Flash-RL model implementation. Inspected at checkpoint revision 5711b268169967567844e1e560e8a3966da959b1. Relevant classes: MiMoV2MoEGate, MiMoV2MoE, MiMoV2Attention, MiMoV2DecoderLayer, vision patch/merger and audio classes. The router's noaux_tc branch explicitly raises when self.training is true. This establishes a limit of this implementation, not a claim that no other training system exists. Pinned source
  • A03. Ainslie et al. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints, 2023. Primary definition and motivation for grouped-query sharing. The chapter does not reproduce the paper's speed/quality measurements as local results. Paper
  • A04. DeepSeek-AI. DeepSeek-V3 Technical Report, first submitted 27 December 2024. Primary evidence for 671B total/37B active, MLA, MoE, 14.8T tokens and MTP. Historical reference, not a latest-release assertion. Paper
  • A05. Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, 2022. Supports distinction between exact IO-aware execution and changing the attention graph. Paper
  • A06. Dao and Gu. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, 2024. Primary Mamba-2/SSD paper. Paper
  • A07. State Spaces. Official Mamba-2 module. Inspected constructor, sequence path, step/cache handling and convolution/state structure. File revision 6b72c122713bb769cc82c6b8e6d019c53d27d6a1, dated 7 October 2024. No kernel was executed. Pinned module
  • A08. Kimi Team. Kimi Linear: An Expressive, Efficient Attention Architecture, 2025. Primary KDA reference; a gated recurrent-memory explanation, not a promise to reproduce the reported throughput. Paper
  • A09. Xie et al. mHC: Manifold-Constrained Hyper-Connections, version 2, 5 January 2026. Supports constrained mixing and Sinkhorn normalization. The two-scalar worked example in the book is original, not an experimental result from the paper. Paper
  • A10. XiaomiMiMo. Official MiMo-V2.6-Flash-RL card and file listing. The card reports 309B total/15B active, while the dynamic Hub tensor summary reports approximately 311B. Card license label is MIT. The listing contains configuration, implementation, technical-report PDF, modality components and weights. Existence and metadata were verified; weight contents were not downloaded or independently counted. Card and pinned files
  • A11. LLM-Core Xiaomi. MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement, technical-report PDF supplied by the official checkpoint repository. Downloaded as a document and text-extracted; architecture sections 2.1-2.4, pretraining/mid-training section 3, router-freezing discussion and post-training overview inspected. Report rounds Flash to 310B and Pro to 1.02T. It reports Flash 48T pretraining tokens (26T text stage +22T multimodal stage) and Pro 30T (27T+3T). The chapter's discussion of the report stays at the architecture/data-stage level. Pinned technical report. Retrieved PDF SHA256: fb81e6e083801b3358f084ed6be953dc23b0d2e434690f4541d5eae03e01e7af
  • A12. XiaomiMiMo. Official MiMo-V2.6-Pro-RL card and configuration. Checkpoint revision 73875d00b30a89ef8cc353a0b60b0e9f9561952d. Publisher architecture and rounded counts; code/config observations confirm layer pattern and head/expert geometry. Card, pinned configuration
  • A13. Z.ai. Official GLM-5.3-Flash card. Publisher evidence for newly trained multimodal base, 320B total/18B active, sparse/linear hybrid, mHC and 30T corpus. Its MIT card label and linked paper are recorded without assuming publication of all training data. The official family repository separately lists FP8 and BF16 variants. Card, family repository
  • A14. XiaomiMiMo. MiMo-V2.6-Flash-RL configuration, revision 5711b268169967567844e1e560e8a3966da959b1. Directly counted 39 local and 9 global layers in hybrid_layer_pattern; inspected dimensions, expert settings, modality configs, max positions and quantization fields. quant_method is fp8 and store_dtype is mxfp4; ignored tensors and mixed-precision components preclude an all-one-dtype byte claim. Pinned configuration
  • A15. XiaomiMiMo. Separate DFlash draft configuration and implementation in the Flash-RL release. Five layers, window 1,024 and block size eight are in the separate config; it is not identified solely by the main model's MTP field. Pinned draft configuration, draft source
  • A16. XiaomiMiMo. MiMo-V2.6-Flash-MOPD card. Verified distinct post-training release and the card's teacher/prefix and repetition-mitigation description. Checkpoint revision 2479e2d0029eca9a34cc7e7f55a121925f81908e; retrieved metadata created 27 September 2026. Its Flash config was inspected alongside RL config. The card also links Pro-MOPD. Card, pinned config
  • A17. GLM-5 Team. GLM-5: from Vibe Coding to Agentic Engineering, arXiv 2602.15763. Verified as the report linked by the current official Flash card. Used only as a family reference, not as sole proof of the later Flash architecture. Paper
  • A18. Z.ai. GLM-5.3-Flash configuration, revision eb9eb208eb0d988989d07a6a12d0fdeb5f52574a. Direct observations: 34 linear/11 sparse-attention layers, 3 dense/42 MoE FFNs, 288 routed experts/top8+1 shared, four mHC streams, KDA geometry, index settings and modality config. indexer_types all read full in this release, even though supporting code allows reuse. Pinned configuration
  • A19. Hugging Face Transformers. modeling_glm5_next.py. Inspected actual KDA sequence/decode paths, FP32 recurrent-state storage, top-k router, mHC mixing, sparse indexer, compressed KV cache and expansion/reference attention path. File revision 0a896aa41bba78f92338db21cc468fe043888657, dated 2 October 2026 at 03:23:05 UTC, before this research. It is a current source inspection, not a tested installed package. Pinned implementation
  • A20. Hugging Face Transformers. GLM-5.3-Flash model documentation. Direct statement that this implementation does not include an MTP layer. The checkpoint's saved transformers_version does not guarantee that any arbitrary installed release supports every component. Documentation source
  • A21. Qwen. Official Qwen3.5-397B-A17B card. 397B/17B, 60-layer 3-to-1 hybrid layout, MoE geometry and native-versus-extended context. Used as a selected architecture comparison rather than a latest-Qwen claim. Card
  • A22. Qwen. Qwen3.5-397B-A17B configuration. Model repository revision returned by metadata: 8472618112abcbd45acbcdc58436aff4233c23f7. Config read directly; layer_types confirms 45 linear and 15 full-attention layers. Pinned configuration
  • A23. NVIDIA. cuSPARSELt data types and supported sparsity formats. Primary source for pattern-specific sparse-kernel contracts. This is not a performance claim for an unspecified consumer RTX. Documentation
  • A24. Rajbhandari et al. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models, 2020 revision. Primary reference for reducing redundancy by sharding training states. The book does not extrapolate the paper's measured scaling to the user's hardware. Paper
  • A25. PyTorch. FSDP/fully-shard documentation and NVIDIA Megatron Core parallelism documentation. Framework references for the distinction between model replicas and sharded parameters, tensor/pipeline/expert placement. PyTorch FSDP, Megatron Core. Verify runtime-specific meanings and support before implementing distributed work; no cluster recipe is promised in the book.

Training Lessons

  • T01. micrograd's scalar engine. Andrej Karpathy, micrograd/engine.py. Inspected the full engine, including graph traversal, backward closures, repeated-use accumulation, and ReLU. Suitable before tensors become a distraction. Educational scalar operations are not a performance model for real training. Revision: 7bc720e951fe422b8f8814aa5aa1b64121d26b4c, default-branch commit 2026-08-03. Repository license: MIT. CPU learning exercise; no consumer-GPU benchmark claimed.
  • T02. micrograd's correctness tests. test/test_engine.py, same revision. Inspected both tests comparing outputs and derivatives against PyTorch, including expressions that reuse values. Teaches an independent reference check and numeric tolerance. Tests were read, not run here; PyTorch was not installed for this review.
  • T03. micrograd video and corrected companion. Karpathy, original YouTube lesson, published 2022-08-16; and second-half lecture notebook. Access: author's YouTube description and chapter list read on the original YouTube page; substantive companion notebook code read. The notebook's exp backward method contains an explicit correction from assignment to accumulation. Useful author chapter markers: 08:08, 1:22:28, 2:01:12. Course repository: karpathy/nn-zero-to-hero, MIT; revision 73c3fcc741f0ec104ca850b1fb0df90e7e8d4cde, 2024-02-20. The course's correct repository name is nn-zero-to-hero, not neural-networks-zero-to-hero. Official syllabus states programming and introductory mathematics prerequisites; this book must teach that missing on-ramp itself.
  • T04. an official visual explanation of gradient descent. Grant Sanderson, Gradient descent, how neural networks learn, lesson dated 2017-10-16; official text adaptation by Josh Pullen. Read the prediction-function/cost-function distinction, local gradient direction, learning-rate explanation, and held-out generalization discussion. Use for intuition after one concrete update. No figures or extended prose from it are reproduced in this book; the official written adaptation was read, not video playback.
  • T05. makemore's MLP lecture notebook. Karpathy, makemore_part2_mlp.ipynb, course revision above. Read all code cells: whole-item train/dev/test partitioning, context construction, embedding lookup, minibatching, cross-entropy, manual updates, and sampling. Original lecture is linked by the creator's course README; its companion code, rather than captions, supplied the substantive evidence.
  • T06. makemore as an executable educational trainer. makemore README. Read data format, scope, model choices, usage, default tiny transformer, and source of the example names. It is a separate executable project from the evolving lecture notebooks. Revision 988aa59e4d8fefa526d06f3b453ad116258398d4, default-branch commit 2022-11-20; MIT. The README is evidence for intended educational/CPU use, not an independently measured runtime.
  • T07. diagnosing activation and gradient scale. makemore_part3_bn.ipynb, course revision above. Inspected training code, running statistics, gradient retention, histogram generation, saturation fraction, and update-to-weight statistics. Useful after a learner has a training loop to debug. The exercise's BatchNorm and tanh examples should not be generalized into a universal transformer architecture recipe.
  • T08. corrected small GPT source for the video. ng-video-lecture/gpt.py. Full file inspected. Its attention divides by the square root of the key/head dimension and masks future positions. It demonstrates a clear pedagogical decoder but is not a production serving or complete checkpointing system. Revision 52201428ed7b46804849dea0b3ccf0de9df1a5c3, 2023-02-07. GitHub returned no license metadata for this repository; no reuse permission is inferred from its public visibility. This book links to it for study and uses independently written code. The final generation call is outside a no-grad/evaluation wrapper, so the book should retain its own explicit inference-mode hygiene rather than copying the script wholesale.
  • T09. GPT video chapter markers and author corrections. Karpathy, original GPT lesson, published 2023-01-17. Read the original description, exercises, chapter list, and corrections on the original YouTube page. No claim of watching the entire video or reading a caption transcript. Useful markers: 14:27 for batches; 47:11 for weighted aggregation; 1:02:00 for learned self-attention; 1:26:48 for residual connections. Author corrections concern future-token masking at 57:00 and head-dimension scaling at 1:20:05. Code and paper are the authority for the corrected operations.
  • T10. attention notation and its visual interpretation. Sanderson, Attention in transformers, step-by-step. Read its query/key/value explanation, notation warning, masking, dimensions, and illustrative-behavior caveat. Its column-oriented diagrams differ from the original paper's row-oriented convention. This is an excellent opportunity to teach shape annotations; attention visualizations should not be mistaken for a complete causal explanation of the trained model.
  • T11. connect the implementation to the original Transformer. Vaswani et al., Attention Is All You Need, HTML v7. Read sections 3.2.1-3.2.3 and 5.2 for scaled attention, projections, masking, and the eight-P100 training premise. Suggested just-in-time question: which operation in the reader's code corresponds to each part of the attention equation? This is a focused paper-reading bridge; the architecture chapter also discusses this paper.
  • T12. inspect a complete small training process. nanoGPT/train.py, revision 3adf61e154c3fe3fca428ad6bc3818b27a3b8291, 2025-11-12; MIT. Inspected accumulation, mixed-precision scaling/clipping order, evaluation, checkpoint dictionary, and logging. Its readable structure is useful for process literacy. Its checkpoint fields do not establish complete deterministic replay; its final-microbatch log is not an exact accumulated-batch average.
  • T13. deprecated examples are historical evidence. nanoGPT README, same revision. Read the deprecation notice, toy-run/reproduction distinction, and original hardware premise. Its GPT-2 reproduction used eight A100 40 GB GPUs for about four days. Include only as a brief caution about transferring old hardware/API assumptions, not as a recommended current trainer or model survey. The small character-model configuration was inspected to confirm it was a distinct experiment.
  • T14. Raschka's minimal training loop and license boundary. Sebastian Raschka, gpt_train.py, requirements, and license. Read the standalone loop, evaluation mode, no-grad handling, token counter, split, small batch/context settings, and dependencies. Revision faa6205602f487c0382295c698ffafe3d635af39, 2026-10-01. The LICENSE.txt is based on Apache 2.0 but modifies the definition of source to exclude books and related images; GitHub metadata returned NOASSERTION. Do not label the entire book/artwork corpus standard Apache-2.0 or reproduce it under the code's terms. Dependency lower bounds are not an exact environment lock.
  • T15. make instruction-loss labels inspectable. gpt_instruction_finetuning.py, same revision. Inspected dataset rendering and custom_collate_fn, especially shifted targets, padding masks, retained terminal token, and truncation. The collator does not mask all instruction positions by default. The transferable lesson is to inspect the actual objective before making an “assistant-only training” claim.
  • T16. LoRA as a loading and export workflow. Hugging Face, smol-course LoRA/PEFT lesson. Read the adapter-loading, merge, and trainer examples. Revision f445ae5d9dd83f355ce6a2b1c6c71b588c8e1541, 2026-09-17; Apache-2.0. A useful bridge after a learner understands which parameters receive gradients. The original LoRA paper was checked as a reading pointer; the LLM chapter supplies the detailed paper discussion.
  • T17. an instructive tutorial-version mismatch. smol-course SFT lesson, hands-on chapter, and root requirements. Inspected actual configuration, formatting, optional adapter, trainer, and upload blocks. The root lock names Torch 2.5.1, Transformers 4.46.3, and TRL 0.12.1; the lesson uses newer model/API examples. The ordinary trainer example's args=config differs from the nearby defined training_config. Its full-fine-tuning snippets are not verified 24 GB recipes.
  • T18. compare against the declared trainer version. TRL 0.12.1 SFT documentation. Read its max_seq_length configuration and completion-only collator discussion. This provides concrete evidence for checking API generation rather than combining contemporary examples with an old requirements file. Its token-context and EOS/padding cautions also show why a collator needs example-level inspection. Use the book's selected trainer version consistently.
  • T19. embedding training as an evaluated retrieval process. Sentence Transformers, NLI example and MS MARCO training example. Full scripts read. Revision 4a3b5cd6ec718e421f57e824a41ed3fd99595df6, 2026-09-21; Apache-2.0. Focus on initial evaluation, negative construction/filtering, duplicate avoidance, cached contrastive minibatches, and post-training evaluation. The MS MARCO script materializes corpus/query/teacher-score dictionaries in host RAM; caching does not eliminate those costs. Both inspected scripts attempt Hub upload despite “optional” comments. Their newer main-branch import layout should not be pasted into the embedding chapter's pinned v5.1.1 environment. Sentence-BERT is the theory bridge already covered in that chapter's source ledger.
  • T20. toy denoising versus actual diffusion. Jonathan Whitaker and Hugging Face, Diffusion Models from Scratch notebook, creator's companion lesson, and embedded original walkthrough. Notebook Markdown/code cells read, including the minimal UNet, uniform corruption, image-prediction objective, and comparison with Gaussian DDPM, noise prediction, timestep conditioning, and sampling. Video link verified as the companion page's embedded lesson; no video/transcript viewing claimed and no timestamp invented. Course revision 57b371aea6d477644726653c3318c3f36afe461c, 2026-09-17; Apache-2.0. The latest commit changes workflow maintenance, not necessarily the 2022-23 teaching content. Unpinned notebook installation cells are not a reproducible lock. The DDPM paper is already part of the diffusion bibliography.
  • T21. follow audio representation through a training loop. Diffusion for Audio notebook, same course revision. Read code/Markdown for resampling, random slicing, spectrogram conversion, batch construction, denoising objective, reconstruction, and upload. Its loss.backward(loss) call is a real source-code pitfall; use ordinary scalar backward unless an upstream-gradient weighting is intentional. Treat its pretrained model and music data as separately licensed inputs, not covered automatically by the course license. The generated model-card template's license: mit is not evidence that all resulting weights can legally receive that label.
  • T22. check the mathematical meaning against framework code. PyTorch 2.8, Tensor.backward, confirms the optional argument is an upstream gradient and gradients accumulate in leaves. Diffusers v0.40.0, DDPM scheduler source, inspected add_noise and prediction-type branches. It uses square-root signal/noise coefficients and distinguishes epsilon, sample, and velocity prediction. These correct possible confusion between variance and standard-deviation notation in informal tutorials. The Audio Diffusion API confirms that this pipeline uses Mel-spectrogram representation and exposes reconstruction settings. It is a main-branch reference, not the book's environment lock.
  • T23. llm.c teaches verification before low-level optimization. Karpathy and contributors, llm.c README and LayerNorm Python reference. Inspected CPU/single-GPU distinctions, reference testing, and the full LayerNorm forward/backward study. Revision f1e2ace651495b74ae22d45d1723443fd00ecd3a, 2025-05-10; MIT. Useful after tensor training is understood. The CPU starter fine-tunes an existing GPT-2 checkpoint for 40 short-context steps; it is not a scratch-pretraining achievement.
  • T24. nanochat gives an end-to-end process, with larger hardware defaults. nanochat README and CPU/MPS script. README and complete small-run script inspected. Revision 92d63d4e8bb4df75c3b71618f31ddde2378b2bcd, 2026-07-03; MIT according to README. End-to-end tokenizer/pretrain/SFT/evaluation/inference structure is a useful advanced map. The speedrun targets eight H100 GPUs; the README requires adjustment below 80 GB. Author-reported price and runtime are not current retail quotes or RTX measurements.
  • T25. current checkpoint state and dependency separation. nanochat checkpoint manager, base training script, and pyproject.toml, revision above. Read checkpoint serialization/loading, resume and data-loader-state plumbing, loop metadata, tokenizer compatibility check, and dependency declarations. Model and optimizer artifacts are separated; optimizer state is rank-specific for distributed runs. The inspected project pins Torch 2.9.1 and selects CPU or CUDA 12.8 wheels through mutually exclusive extras. This is its own environment, not a reason to mix its dependencies into the book’s other projects. No deterministic cross-device resume or RTX fit claim is inferred.