Works cited
Checked 2 October 2026 unless a source entry states otherwise. A cited implementation can explain a design without being a measured reproduction. Repository main branches can change; use the pinned versions and captured revisions in each project.
Foundation
- F01. PyTorch contributors. Autograd mechanics. Official reference on reverse automatic differentiation and computation graphs. https://docs.pytorch.org/docs/stable/notes/autograd.html
- F02. PyTorch contributors. AdamW reference, PyTorch 2.8. Parameter meanings and update algorithm. https://docs.pytorch.org/docs/2.8/generated/torch.optim.AdamW.html
- F03. Vaswani, A. and colleagues. Attention Is All You Need. 2017. The original Transformer architecture paper. https://arxiv.org/abs/1706.03762
- F04. Hugging Face. Model training anatomy, Transformers 4.48.0. A documented example of mixed-precision training-state accounting, not a universal bytes-per-parameter law. https://huggingface.co/docs/transformers/v4.48.0/model_memory_anatomy
- F05. PyTorch contributors. torch.utils.checkpoint. Activation checkpointing and recomputation behavior. https://docs.pytorch.org/docs/stable/checkpoint.html
- F06. PyTorch contributors. torch.cuda.max_memory_allocated. Peak tensor-allocation measurement. https://docs.pytorch.org/docs/stable/generated/torch.cuda.max_memory_allocated.html
- F07. PyTorch contributors. CUDA semantics. Allocated versus reserved memory and the limits of empty_cache. https://docs.pytorch.org/docs/stable/notes/cuda.html
- F08. Hoffmann, J. and colleagues. Training Compute-Optimal Large Language Models. 2022. Empirical scaling of parameters and data under compute budgets; not a universal prescription for every domain. https://arxiv.org/abs/2203.15556
- F09. scikit-learn developers. Common pitfalls and recommended practices. Train/test separation and preprocessing leakage. https://scikit-learn.org/1.9/common_pitfalls.html
- F10. PyTorch Foundation. PyTorch 2.8 Release Blog. 6 August 2025. Verified release used as the tiny-transformer reference API pin. https://pytorch.org/blog/pytorch-2-8/
- F11. PyTorch contributors. scaled_dot_product_attention, PyTorch 2.8. Causal masking, dropout behavior, and supported attention implementations. https://docs.pytorch.org/docs/2.8/generated/torch.nn.functional.scaled_dot_product_attention.html
- F12. PyTorch contributors. Automatic Mixed Precision examples. Gradient scaling, clipping order, and accumulation. https://docs.pytorch.org/docs/stable/notes/amp_examples.html
- F13. PyTorch contributors. CrossEntropyLoss. Logit targets, reduction, and ignore_index. https://docs.pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html
- F14. PyTorch contributors. Reproducibility. Seeds, deterministic operations, and cross-platform limitations. https://docs.pytorch.org/docs/stable/notes/randomness.html
- F15. PyTorch Foundation. Previous PyTorch Versions. Official wheel commands for CPU and CUDA variants, including torch 2.8.0 with torchvision 0.23.0 and torch 2.7.1 with torchvision 0.22.1. https://pytorch.org/get-started/previous-versions/
Tiny Ml Embedding
- M01. scikit-learn 1.8, Common pitfalls and recommended practices. Supports training-only preprocessing, pipelines, and consistent inference transformations
- M02. scikit-learn 1.8, Probability
calibration. Supports reliability interpretation, independent
calibration data,
FrozenEstimator, sigmoid calibration, and limitations of Brier/log-loss interpretation
- M03. scikit-learn 1.8, HistGradientBoostingClassifier, and LogisticRegression. Exact estimator controls for the CPU classification example
- M04. scikit-learn 1.8, HistGradientBoostingRegressor. Exact tree-regression controls
- M05. PyTorch 2.8, BCEWithLogitsLoss and CrossEntropyLoss. Logit/target conventions and stable loss APIs; the teaching Dice term is implemented explicitly in the accompanying script
- M06. torchvision 0.23, MobileNetV3-Small. Explicit weight enum and associated transforms
- M07. Ronneberger, Fischer, and Brox, U-Net: Convolutional Networks for Biomedical Image Segmentation, 2015. Encoder-decoder/skip-connection source; the supplied tiny network is a pedagogical variation, not a reproduction claim
- M08. Cheng et al., Boundary IoU: Improving Object-Centric Image Segmentation Evaluation, 2021; original CVPR paper. Motivation for evaluating boundaries separately from region overlap. The included boundary F1 is a different, explicitly specified metric
- M09. OpenCV, Structural
Analysis and Shape Descriptors.
findContours, hierarchy/retrieval modes, and polygon-approximation semantics. This older versioned reference documents the classic API used by the optional contour recipe
- M10. PyTorch, TorchVision Object Detection Finetuning Tutorial. Dataset contract, box/instance masks, label conventions, and pretrained Mask R-CNN fine-tuning. The tutorial states torchvision >=0.16 for its current variant; follow the pinned 0.23 API when combining with this book
- M11. torchvision BSD-3-Clause source license and README's separate dataset/model-weight licensing warnings. Permissive code does not settle training-data or pretrained-weight rights
- M12. PyTorch 2.8, Automatic Mixed
Precision.
torch.autocast,torch.amp.GradScaler, and deprecation of old AMP namespaces
- M14. Sentence Transformers v5.1.1, MultipleNegativesRankingLoss
source. Exact
model,scale, similarity, and forward-call behavior used in the example
- M15. Sentence Transformers, Samplers. Duplicate-avoidance rationale; the book's direct loop performs its own one-document-per-batch grouping
- M16. Gao et al., Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup, 2021, and Sentence Transformers loss reference. Gradient-cache tradeoffs and cached contrastive options; no claim that all memory costs vanish
- M17. Radford et al., Learning Transferable Visual Models From Natural Language Supervision, 2021; arXiv version. Primary CLIP dual-encoder/multimodal contrastive example; no claim to reproduce its large-scale pretraining on one GPU
- M18. Sentence Transformers, Evaluation reference. Information-retrieval evaluator terminology and supported metrics; this book includes a small independent exact-search metric implementation
- M19. PyTorch, ExecuTorch project and How ExecuTorch works. Export, optimization, and target-runtime deployment workflow; no tested mobile artifact is supplied
- M20. scikit-learn, Cross-validation guide. Group-aware and time-aware evaluation principles; this source was the stable guide at verification, while the executed estimator version is 1.8.0
- M21. torchvision, official compatibility table. PyTorch 2.8 pairs with torchvision 0.23 and supports Python 3.9-3.13 according to the published table
- M22. Reimers and Gurevych, Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, 2019. Primary background reading on independently encoded sentence representations
- M23. Schroff, Kalenichenko, and Philbin, FaceNet: A Unified Embedding for Face Recognition and Clustering, 2015. Primary triplet-learning background; the book discusses the objective, not a biometric deployment recommendation
- M24. scikit-learn, Model persistence. Security and compatibility cautions for pickle-derived model artifacts
- M25. torchvision 0.23, Faster R-CNN MobileNetV3-Large 320 FPN. Official architecture/weight interface; use a custom target dataset and review its rights separately
- M26. Sentence Transformers, v5.1.1 release. The named version exists; this is an API target, not a tested environment lock
- M28. torchvision v0.23.0, official detection training process and reference scripts. Published scripts, losses/evaluation support, and documented original training commands. Many original commands use eight GPUs; they are process references, not recipes promised to fit one 24 GB GPU unchanged
- M29. Sentence Transformers, official embedding training examples. Public examples for retrieval, similarity, triplet, and other supervision styles; main-branch code can differ from the pinned v5.1.1 direct-loop recipe
Current Embedding
- M30. Qwen, Qwen3-Embedding-0.6B official model card. Public/non-gated at verification, Apache-2.0 metadata, 0.6B family label, 1024-dimensional output, 32K stated context, minimum library requirements, and recommended query-only instructions. No leaderboard claim is used as an application-quality guarantee
- M31. Qwen, pinned pooling configuration and released module graph. Last-token pooling, no mean pooling, 1024 features, and normalization module. Sentence Transformers 5.1.1 Pooling implementation shows last non-padding-token selection
- M32. Qwen, official prompt configuration, pinned tokenizer configuration, and model repository file history. Query instruction/document empty-prefix distinction and native tokenizer artifacts; the lab supplies its own domain instruction in the documented format consistently across all paths
- M33. Qwen, official architecture configuration and Transformers 4.57.1 Qwen3 implementation. Shape-based parameter calculation and Qwen3 backbone configuration. Derived count is 595,776,512 for this backbone, not a claim of author-executed loading; the program reports the actual loaded count
- M34. SentenceTransformer v5.1.1 implementation, MultipleNegativesRankingLoss v5.1.1 implementation, Qwen3 v4.57.1 implementation, and PyTorch 2.8 AMP API. Source-reviewed loader/tokenizer/forward/loss, SDPA-compatible model, and autocast APIs. This is not an empirically tested dependency lock
- M35. PyTorch, official previous-version installation commands, v2.8.0. CPU/CUDA 12.8 wheel indexes and matching pinned release. Hardware/driver compatibility must still be established locally
- M36. Qwen, verified immutable model revision 97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3 and repository files/sizes. Source identity and approximate BF16-weight download size
- M37. Qwen, official Qwen3 Embedding repository and Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. Current public model/process and original research references. The lab is a small full-fine-tuning example, not a reproduction of Qwen's original large-scale training
Llm Training
- L01. Transformers chat templates. URL: https://huggingface.co/docs/transformers/chat_templating Inspected: Current documentation; role serialization, apply_chat_template, training/inference generation-prompt and duplicate-special-token cautions
- L02. Qwen3-0.6B official model card. URL: https://huggingface.co/Qwen/Qwen3-0.6B Inspected: Model Overview; Quickstart; thinking/non-thinking mode; Apache-2.0. Card says 0.6B, 28 layers, 16 Q/8 KV heads and 32,768 context. No quality benchmark copied
- L03. Qwen3-0.6B official Hub metadata. URL: https://huggingface.co/api/models/Qwen/Qwen3-0.6B Inspected: Inspected sha c1899de289a04d12100db370d81485cdf75e47ca, tokenizer template, 751,632,384 serialized BF16 entries. This is not a unique trainable-parameter count
- L04. Transformers Trainer tagged implementation. URL: https://raw.githubusercontent.com/huggingface/transformers/v5.18.0/src/transformers/trainer.py Inspected: v5.18.0; constructor processing_class and optimizer_cls_and_kwargs; causal-LM shifted target counting; quantized model placement; loss-only evaluation design
- L05. TRL supervised fine tuning trainer. URL: https://huggingface.co/docs/trl/v1.14.1/sft_trainer Inspected: v1.14.1; expected dataset formats; completion/assistant masks; training templates; tool columns; default chunked_nll. Used for API review, not a copied runnable benchmark
- L06. TRL SFTConfig tagged implementation. URL: https://raw.githubusercontent.com/huggingface/trl/v1.14.1/trl/trainer/sft_config.py Inspected: v1.14.1; max_length, assistant_only_loss, completion_only_loss, padding/packing fields and loss_type
- L07. Transformers releases. URL: https://github.com/huggingface/transformers/releases Inspected: Observed v5.18.0 release; tagged source inspected separately. Release existence does not prove installed compatibility
- L08. Accelerate releases. URL: https://github.com/huggingface/accelerate/releases Inspected: Observed v1.15.0 release and release notes; used only to identify review pin
- L09. PEFT releases. URL: https://github.com/huggingface/peft/releases Inspected: Observed v0.21.2 release, including Transformers 5.18 compatibility fix described for encoder-decoder models
- L10. bitsandbytes releases. URL: https://github.com/bitsandbytes-foundation/bitsandbytes/releases Inspected: Observed stable 0.50.2, distinguished from continuous prerelease wheel
- L11. TRL releases. URL: https://github.com/huggingface/trl/releases Inspected: Observed v1.14.1; reviewed tagged implementation and metadata rather than relying on latest docs alone
- L12. Datasets releases. URL: https://github.com/huggingface/datasets/releases Inspected: Observed 5.0.1 release; optional TRL environment pin
- L13. Official previous PyTorch versions. URL: https://pytorch.org/get-started/previous-versions/ Inspected: Inspected v2.12.1 CUDA 12.6 wheel command; this is a chosen reproducible release, not a claim that it is newest. No installation performed
- L14. Hu et al LoRA. URL: https://arxiv.org/html/2106.09685v2 Inspected: arXiv v2, 2021; low-rank reparameterization and frozen pretrained weights. Matrix example in book calculated independently
- L15. PEFT LoRA API reference. URL: https://huggingface.co/docs/peft/v0.21.0/package_reference/lora Inspected: Official docs resolved to v0.21.0; LoraConfig rank, alpha, target_modules, all-linear behavior. Code pins patch release 0.21.2
- L16. PyTorch AdamW API. URL: https://docs.pytorch.org/docs/2.12/generated/torch.optim.AdamW.html Inspected: PyTorch 2.12; moment updates, optimizer options and foreach memory caveat. Book memory rows state their own precision assumptions
- L17. PyTorch automatic mixed precision. URL: https://docs.pytorch.org/docs/2.12/amp.html Inspected: PyTorch 2.12; autocast changes selected compute precision, distinct from stored parameter dtype
- L18. Dettmers et al QLoRA. URL: https://arxiv.org/html/2305.14314v1 Inspected: arXiv v1, 2023; frozen four-bit base, NF4, double quantization and paged optimizers. No paper hardware result represented as a local measurement
- L19. PEFT quantization guide. URL: https://huggingface.co/docs/peft/v0.21.0/developer_guides/quantization Inspected: v0.21.0; BitsAndBytesConfig preparation and prepare_model_for_kbit_training before attaching adapters
- L20. Transformers bitsandbytes integration. URL: https://huggingface.co/docs/transformers/quantization/bitsandbytes Inspected: Current official docs; NF4, compute dtype, double quantization, extra-parameter training restriction and memory footprint
- L21. Lewis et al Retrieval Augmented Generation. URL: https://arxiv.org/abs/2005.11401 Inspected: Abstract and publication metadata; parametric generation combined with non-parametric retrieval memory
- L22. Gururangan et al Do not Stop Pretraining. URL: https://arxiv.org/abs/2004.10964 Inspected: Abstract and publication metadata; domain/task adaptation in studied settings. Not treated as a universal modern-decoder guarantee
- L23. Meng et al Locating and Editing Factual Associations in GPT. URL: https://arxiv.org/abs/2202.05262 Inspected: Abstract and metadata; ROME factual-association editing scope. No unsupported empirical locality rate used
- L24. Meng et al Mass Editing Memory in a Transformer. URL: https://arxiv.org/abs/2210.07229 Inspected: Abstract and metadata; MEMIT many-edit research scope. Engineering cautions in book are recommendations, not claimed study results
- L25. Transformers tool use templates. URL: https://huggingface.co/docs/transformers/chat_extras Inspected: Current official docs; structured schemas, tool_calls and tool-role observations, model-specific serialization
- L26. TRL dataset formats and types. URL: https://huggingface.co/docs/trl/dataset_formats Inspected: Current official docs; conversational, preference and tool-calling structures. Local fixtures are original fictional data
- L27. Berkeley Function Calling Leaderboard. URL: https://gorilla.cs.berkeley.edu/leaderboard.html Inspected: Inspected V4 methodology links; snapshot says models evaluated at f7cf735 and bfcl-eval==2025.12.17. No leaderboard number reported as an author measurement
- L28. BFCL multi-turn methodology. URL: https://gorilla.cs.berkeley.edu/blogs/13_bfcl_v3_multi_turn.html Inspected: Official BFCL-v3 multi-turn/multi-step description; used to motivate separation of call syntax and interaction outcomes
- L29. Rafailov et al Direct Preference Optimization. URL: https://arxiv.org/abs/2305.18290 Inspected: arXiv v3 metadata and abstract; reference-relative pairwise objective checked against TRL math in L30
- L30. TRL DPO trainer. URL: https://huggingface.co/docs/trl/v1.14.1/dpo_trainer Inspected: v1.14.1; prompt/chosen/rejected records, equation, reference model and precomputed reference log probabilities
- L31. Hinton et al Distilling the Knowledge in a Neural Network. URL: https://arxiv.org/abs/1503.02531 Inspected: Abstract and metadata; teacher distributions and distillation origin. Temperature explanation is standard derivation, not a reproduced result
- L32. Kim and Rush Sequence Level Knowledge Distillation. URL: https://arxiv.org/abs/1606.07947 Inspected: Abstract and metadata; sequence-level teacher targets distinguished from logit matching
- L33. TRL Distillation Trainer. URL: https://huggingface.co/docs/trl/distillation_trainer Inspected: Current official docs; DistillationTrainer and on/off-policy modes. No untested trainer code presented as a measured run
- L34. Qwen3-0.6B official config. URL: https://huggingface.co/Qwen/Qwen3-0.6B/blob/main/config.json Inspected: Inspected config at main, last config change shown 167b810; head_dim=128, hidden=1024, intermediate=3072, 28 layers, vocab=151936, tied embeddings. Full model pinned via L03
- L35. Qwen3-1.7B official model card. URL: https://huggingface.co/Qwen/Qwen3-1.7B Inspected: Model Overview and license; nominal 1.7B, 28 layers, 16 Q/8 KV GQA heads and 32,768 published context
- L36. Qwen3.5-0.8B official model card. URL: https://huggingface.co/Qwen/Qwen3.5-0.8B Inspected: Model Overview; LM parameters separate from vision encoder, 24 layers, DeltaNet/gated attention layout, padded vocabulary 248320; Apache-2.0
- L37. Qwen3.5-2B official model card. URL: https://huggingface.co/Qwen/Qwen3.5-2B Inspected: Model Overview; 2B LM label, vision encoder, 24 layers, hybrid layout. No benchmark transferred to local hardware
- L38. Google Gemma 3 270M release article. URL: https://developers.googleblog.com/introducing-gemma-3-270m/ Inspected: Core capabilities; 270M total, 170M embedding/100M transformer parameters and task-specific intent. Phone-energy result deliberately not generalized
- L39. Gemma 3 1B instruction model card. URL: https://huggingface.co/google/gemma-3-1b-it Inspected: Model Information, access conditions, context, license label. Family-level multimodal text not applied indiscriminately to 1B
- L40. Transformers Gemma 3 documentation. URL: https://huggingface.co/docs/transformers/model_doc/gemma3 Inspected: Architecture and text-model class documentation; small text-only path distinguished from larger vision-language models
- L41. Granite 4.0 350M official model card. URL: https://huggingface.co/ibm-granite/granite-4.0-350m Inspected: Model Architecture table; dense baseline, 28 attention layers, GQA, 350M; Apache-2.0. Card metadata tag is not used as architecture authority
- L42. Granite 4.0 H 350M official model card. URL: https://huggingface.co/ibm-granite/granite-4.0-h-350m Inspected: Model Architecture; 340M total hybrid, attention and Mamba2. No claim of a tested CUDA kernel path
- L43. Granite 4.0 1B official model card. URL: https://huggingface.co/ibm-granite/granite-4.0-1b Inspected: Architecture table gives 1.6B dense and 1.5B hybrid, despite shortened names. Used to flag strict parameter ceilings
- L44. Granite 3.3 2B instruction model card. URL: https://huggingface.co/ibm-granite/granite-3.3-2b-instruct Inspected: Model Summary, Apache-2.0 and intended domain/tool/RAG use. Not described as latest Granite
- L45. Granite 3.3 2B official config. URL: https://huggingface.co/ibm-granite/granite-3.3-2b-instruct/blob/main/config.json Inspected: Inspected 40 layers, hidden=2048, intermediate=8192, Q=32/KV=8, vocab=49159 and tied embeddings
- L46. SmolLM2 360M instruction model card. URL: https://huggingface.co/HuggingFaceTB/SmolLM2-360M-Instruct Inspected: Official model card and training-resource links; small checkpoint candidate
- L47. Meta Llama 3.2 1B instruction model card. URL: https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct Inspected: Model information, GQA and license label. License/gating require checkpoint-specific review; no claim of downloaded gated weights
- L48. Granite 4.2 3B official model card. URL: https://huggingface.co/ibm-granite/granite-4.2-3b Inspected: Model Summary/Design; August 25 2026, dense reasoning model and nominal 3B. Exact total cap not asserted
- L49. SmolLM3 3B official model card. URL: https://huggingface.co/HuggingFaceTB/SmolLM3-3B Inspected: Training/software/hardware sections; 11T pretraining tokens and 384 H100. Used as published process scale, not practical full-training project
- L50. Qwen3 30B A3B official model card. URL: https://huggingface.co/Qwen/Qwen3-30B-A3B Inspected: Model Overview; 30.5B total, 3.3B activated, 128 experts, 8 activated. Excluded from practical scope
- L51. Gemma 4 E2B instruction model card. URL: https://huggingface.co/google/gemma-4-E2B-it Inspected: Dense-model table and license; 2.3B effective, 5.1B with embeddings, per-layer embeddings, Apache-2.0. Excluded from strict total scope
- L52. SmolLM published fine-tuning script. URL: https://raw.githubusercontent.com/huggingface/smollm/main/text/finetuning/train.py Inspected: Main as inspected; args/defaults and SFTConfig call. Verified max_seq_length, default push_to_hub=True and report_to=wandb. No immutable repository SHA was resolved; do not label this source pinned
- L53. SmolLM public pretraining process. URL: https://github.com/huggingface/smollm/blob/main/text/pretraining/README.md Inspected: Main as inspected; launch example, intra-document masks, configuration paragraph, 2.36M global-token batch, 384 H100 for 24 days. Linked logs not re-run
- L54. Hoffmann et al Training Compute Optimal Large Language Models. URL: https://arxiv.org/html/2203.15556v1 Inspected: arXiv v1; section 3 and efficient-frontier 6ND approximation, tables/context. Book 1B/20B-token/50TFLOP example is explicitly hypothetical
- L55. LFM2 350M official model card. URL: https://huggingface.co/LiquidAI/LFM2-350M Inspected: Model details table with exact counts; hybrid blocks, custom license, tool format, link to newer LFM2.5-350M
- L56. LFM2.5 350M official model card. URL: https://huggingface.co/LiquidAI/LFM2.5-350M Inspected: Model Details, Tool Use, Fine-Tuning, exported formats; 16 blocks, 32K context, vocab=65536, custom LFM1.0 label. Vendor speed claims not reused as local measurements
- L57. Liquid official TRL fine-tuning guide. URL: https://docs.liquid.ai/lfm/fine-tuning/trl Inspected: LoRA Fine-Tuning and full-update examples; attention projections explicitly targeted; inspected old tokenizer keyword and minimum-version install string. Current API correction comes from L05
- L58. Transformers LFM2 model documentation. URL: https://huggingface.co/docs/transformers/model_doc/lfm2 Inspected: Native Lfm2ForCausalLM with input_ids, labels, attention_mask and cache arguments; proves supported model interface, not all-backend training success
- L59. Liquid official llama.cpp deployment guide. URL: https://docs.liquid.ai/deployment/on-device/llama-cpp Inspected: GGUF downloading, CPU-first execution, CLI/server and GPU-offload distinctions; model-specific existing export route, not proof of arbitrary adapter conversion
- L60. LFM2 technical report. URL: https://arxiv.org/abs/2511.23404 Inspected: arXiv v1 abstract/metadata; hybrid gated-short-convolution/GQA design and hardware-in-the-loop search. No claimed reproduction of reported CPU gains
Jev Architecture
- J01. TypeSafe's announcement. https://typesafe.ai/blog/introducing-system-one-models-and-jev
- J02. official API. https://docs.typesafe.ai/api
- J03. Kev. https://github.com/jaredpalmer/kev
- J04. architecture hypothesis and its caveats. https://archerhume.com/posts/jevs-architecture-unmasked
- J05. CLM. https://github.com/Contrastive-LM/CLM
- J06. Official configuration. https://huggingface.co/Qwen/Qwen3-0.6B/blob/main/config.json
- J07. RoFormer paper. https://arxiv.org/abs/2104.09864
- J08. Transformer paper. https://arxiv.org/abs/1706.03762
- J09. GQA paper. https://arxiv.org/abs/2305.13245
- J10. Official implementation. https://github.com/huggingface/transformers/blob/main/src/transformers/models/qwen3/modeling_qwen3.py
- J11. GLU variants paper. https://arxiv.org/abs/2002.05202
- J12. RMSNorm paper. https://arxiv.org/abs/1910.07467
- J13. Kev model source. https://github.com/jaredpalmer/kev/blob/main/kev/model.py
- J14. Official Qwen3.5 configuration. https://huggingface.co/Qwen/Qwen3.5-0.8B-Base/blob/main/config.json
- J15. Kev parity tests. https://github.com/jaredpalmer/kev/blob/main/tests/test_model.py
- J16. KV-cache explanation. https://huggingface.co/docs/transformers/main/en/llm_tutorial_optimization
- J17. Temperature-scaling paper. https://arxiv.org/abs/1706.04599
- J18. CLM head implementation. https://github.com/Contrastive-LM/CLM/blob/main/src/clm/heads.py
- J19. CLM scoring and cache implementation. https://github.com/Contrastive-LM/CLM/blob/main/src/clm/engine.py
- J20. Contrastive predictive coding and InfoNCE. https://arxiv.org/abs/1807.03748
- J21. CLM training implementation. https://github.com/Contrastive-LM/CLM/blob/main/train/finetune.py
- J22. Granite 4.2-3B model card. https://huggingface.co/ibm-granite/granite-4.2-3b
- J23. Kev-0.8B model card. https://huggingface.co/jaredpalmer/kev-0.8b
- J24. Laya official model card. https://huggingface.co/convaiinnovations/laya
- J25. Laya original repository. https://github.com/NandhaKishorM/laya
- J26. Laya shipped model implementation. https://huggingface.co/convaiinnovations/laya/blob/main/rl_common.py
- J27. Laya shipped configuration. https://huggingface.co/convaiinnovations/laya/blob/main/rl_agent_config.json
- J28. Laya single-process MPS trainer. https://github.com/NandhaKishorM/laya/blob/main/notebooks/laya_finetune_typed_decisions_mps.py
- J29. Laya Kaggle 2xT4 notebook. https://github.com/NandhaKishorM/laya/blob/main/notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb
- J30. Laya measured browser-agent adaptation. https://github.com/NandhaKishorM/laya/blob/main/docs/finetune_browser_agent.md
- J31. ModernBERT paper. https://arxiv.org/abs/2412.13663
- J32. ModernBERT-large configuration. https://huggingface.co/answerdotai/ModernBERT-large/blob/main/config.json
- J33. Ten Levels of Jev original repository. https://github.com/disler/ten-levels-of-jev
- J34. Ten Levels provider client. https://github.com/disler/ten-levels-of-jev/blob/main/apps/ten-levels/src/core/client.ts
- J35. Ten Levels confidence thresholds. https://github.com/disler/ten-levels-of-jev/blob/main/apps/ten-levels/src/levels/level04/confidence.ts
- J36. Ten Levels application README. https://github.com/disler/ten-levels-of-jev/blob/main/apps/ten-levels/README.md
- J37. Ten Levels of Jev video. https://www.youtube.com/watch?v=_U-O5lYhJ7Q
- J38. Karpathy: build GPT in code. https://www.youtube.com/watch?v=kCc8FmEb1nY
- J39. nanoGPT source (historical, now deprecated). https://github.com/karpathy/nanoGPT
Diffusion Training
- D01. Ho, Jain and Abbeel, Denoising Diffusion Probabilistic Models (2020). https://arxiv.org/abs/2006.11239 - foundational Gaussian forward corruption, noise-prediction training and reverse generation. Used for the short objective explanation; our waveform code is an original teaching implementation, not a reproduction of their benchmark results.
- D02. Song et al., Score-Based Generative Modeling through Stochastic Differential Equations (2020/2021). https://arxiv.org/abs/2011.13456 - score interpretation and reverse SDE/probability-flow viewpoint. No paper benchmark is presented as a result of our scripts.
- D03. Lipman et al., Flow Matching for Generative Modeling (2022/2023). https://arxiv.org/abs/2210.02747 - vector-field regression along probability paths; distinguishes flow matching from a particular epsilon-prediction recipe.
- D04. Ho and Salimans, Classifier-Free Diffusion Guidance (2022). https://arxiv.org/abs/2207.12598 - conditional/unconditional predictions and inference-time quality/diversity trade-off. The chapter clearly states its guidance-scale convention.
- D05. Rombach et al., High-Resolution Image Synthesis with Latent Diffusion Models (2021/2022). https://arxiv.org/abs/2112.10752 - separates an autoencoder representation from conditional latent diffusion.
- D06. Peebles and Xie, Scalable Diffusion Models with Transformers (2022/2023). https://arxiv.org/abs/2212.09748 - transformer backbone on latent patches; architecture is distinct from output target.
- D07. Ruiz et al., DreamBooth (2022). https://arxiv.org/abs/2208.12242 - subject-driven personalization and class-specific prior preservation; not an alternative mathematical definition of LoRA.
- D17. Nichol and Dhariwal, Improved Denoising Diffusion Probabilistic Models (2021). https://arxiv.org/abs/2102.09672 - source for the cosine cumulative-noise schedule idea. The supplied code fixes s=0.008 and clips beta at 0.999; it does not implement every improvement in the paper.
- D18. Kong et al., DiffWave: A Versatile Diffusion Model for Audio Synthesis (2020/2021). https://arxiv.org/abs/2009.09761 - waveform diffusion and conditional/unconditional audio synthesis. Our U-Net architecture is explicitly not described as DiffWave.
- D08. Hugging Face Diffusers, LoRA training guide. https://huggingface.co/docs/diffusers/training/lora - general parameter-efficient adaptation integration. The current project follows the Sana-specific code below, not this page's legacy SD1.5 command.
- D16. Hugging Face Datasets v3.6.0, Create an image
dataset. https://huggingface.co/docs/datasets/v3.6.0/en/image_dataset
- imagefolder and JSONL metadata relationship. Our extra
groupfield and split discipline are deliberate project design, not a feature that automatically prevents leakage in the upstream trainer.
- D19. Stable Audio Open 1.0 official card. https://huggingface.co/stabilityai/stable-audio-open-1.0 - 47-second, 44.1 kHz stereo limit; autoencoder, T5 and latent DiT components; library usage. Research paper: https://arxiv.org/abs/2407.14358 . These are model characteristics, not a reproduced 24 GB fine-tuning result.
- D20. Stable Audio Open Small official card. https://huggingface.co/stabilityai/stable-audio-open-small - 11-second, 44.1 kHz stereo model and distilled inference recipe. Related paper: https://arxiv.org/abs/2505.08175 . The book does not invent a generic LoRA training flag for this checkpoint.
- D21. Stability AI stable-audio-tools official repository. https://github.com/Stability-AI/stable-audio-tools - separate model/dataset configuration and training machinery. Current README: https://raw.githubusercontent.com/Stability-AI/stable-audio-tools/main/README.md . Presence of general training code is not a validated memory estimate or proof of a particular distilled model's adapter support.
- D22. Stability AI Stable Audio 3 official repository. https://github.com/Stability-AI/stable-audio-3 - model variants, fine-tuning support and published inference performance. README inspected: https://raw.githubusercontent.com/Stability-AI/stable-audio-3/main/README.md . The VRAM table labels H200 inference with unchunked decoding; it is not a consumer-GPU training table. Technical report: https://arxiv.org/abs/2605.17991 .
- D23. Stable Audio 3 official LoRA guide. https://github.com/Stability-AI/stable-audio-3/blob/main/docs/workflows/lora.md - confirms adapter workflow. Approximate memory table and a separate reduced-memory example are insufficiently uniform to treat as reproduced peak training measurements. No quoted training throughput is claimed.
- D24. Stable Audio 3 actual training script. https://github.com/Stability-AI/stable-audio-3/blob/main/scripts/train_lora.py
- inspected parser and training code, including base-model-only loading;
raw audio plus
.txtmetadata; duration, rank, batch, local CSV logging; and demo sample-size handling.--save_diris the actual parser name at the inspected source; do not blindly copy prose mentioning--output_dir. Its full dependency stack was not installed or executed for the book.
- D25. Stable Audio 3 dependency specification. https://github.com/Stability-AI/stable-audio-3/blob/main/pyproject.toml - pins torch/torchaudio 2.7.1 and has its own modern Transformers dependencies. Avoid mixing this environment with the image project's pins.
- D26. Stable Audio 3 Small SFX official model card.
https://huggingface.co/stabilityai/stable-audio-3-small-sfx
- describes the family, base relationship, T5Gemma conditioning and
additional Gemma terms. Note that the card's lower-level sample has a
CPU-path variable issue (
model_halfis not initialized in the CPU branch); the book does not copy that snippet or claim parity with the higher-level API.
- D27. PyTorch previous versions, v2.7.1. https://pytorch.org/get-started/previous-versions/ - official torch 2.7.1 / torchvision 0.22.1 pairing and CUDA 12.6 wheel index. Driver compatibility and actual device availability remain machine-specific.
- D28. Original DDPM implementation. https://github.com/hojonathanho/diffusion - primary author repository. Its README specifies TensorFlow 1.15, Python 3.5 and TPU experiments. Read to connect the paper with its historical implementation; do not present its installation as a modern single-RTX recipe.
- D29. DiffWave reference implementation. https://github.com/lmnt-com/diffwave - public waveform/vocoder training and inference implementation. Useful next reading after the toy class-conditioned example; a mel-conditioned vocoder's data contract is different from text-conditioned sound generation.
- D30. Official Sana 1.6B BF16 model card. https://huggingface.co/Efficient-Large-Model/Sana_1600M_1024px_BF16_diffusers - exact primary checkpoint, 1.648B image transformer, Gemma2-2B-IT text encoding, 32× DC-AE compression, component-specific dtypes, Apache/Gemma terms and research-oriented intended use. The complete pipeline is explicitly larger than the transformer alone.
- D31. Actual Diffusers v0.36.0 Sana LoRA trainer. https://github.com/huggingface/diffusers/blob/v0.36.0/examples/dreambooth/train_dreambooth_lora_sana.py
- inspected argument parser, local caption-dataset path, mixed component
precision, offload, flow target
noise - model_input, checkpoint hooks and resume behavior. Raw inspected source: https://raw.githubusercontent.com/huggingface/diffusers/v0.36.0/examples/dreambooth/train_dreambooth_lora_sana.py . Important code findings: cache-by-batch-position plus shuffled loader can mismatch per-image captions; our command disables that cache. Final pipeline loading is unconditional in upstream even without requested validation; our guarded patch skips that unused allocation after adapter save. Resume restores training state but does not exactly skip to the old next minibatch.
- D32. Actual Diffusers v0.36.0 Sana inference
pipeline. https://github.com/huggingface/diffusers/blob/v0.36.0/src/diffusers/pipelines/sana/pipeline_sana.py
- verified
SanaPipeline, LoRA loading mixin, CPU-offload sequence, prompt embeddings and masks, 32-divisible dimensions, sampling inputs and decoder postprocessing. Training and inference explicitly agree on max_sequence_length=128 and complex_human_instruction=None instead of inheriting the pipeline's different instruction-prefix default.
- D33. Chen et al., SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers (2024/ICLR 2025). https://arxiv.org/abs/2410.10629 - efficient architecture and representation design. Its reported inference hardware/result is not represented as a measured fine-tuning result for the book.
- D34. Diffusers v0.36.0 dependency table. https://github.com/huggingface/diffusers/blob/v0.36.0/src/diffusers/dependency_versions_table.py - checked PEFT/Hub/Transformers requirements against the explicit candidate pins. Actual installation/import tests remain required.
- D35. Official maintained NVlabs Sana repository. https://github.com/NVlabs/Sana - current family, training/inference documentation and model variants. It is actively maintained beyond the original paper; the book selects a stable checkpoint with an explicit fine-tuning route rather than claiming the newest family release automatically has the same trainer.
- D36. Official Diffusers Sana DreamBooth/LoRA guide. https://github.com/huggingface/diffusers/blob/main/examples/dreambooth/README_sana.md - expressly targets the selected BF16 1.6B checkpoint, shows training, and documents offload/cache/optimizer options. The runnable book script is pinned to v0.36.0, not this moving document. No exact consumer-24GB training peak is stated or inferred.
- D37. Official Sana 600M 512-pixel card. https://huggingface.co/Efficient-Large-Model/Sana_600M_512px_diffusers - maintained 590M-transformer alternative, distinct FP16 variant, Gemma and DC-AE companions. It is a possible later experiment, not a drop-in equivalent of the BF16 project or a claim that the full pipeline has only 590M parameters.
Cohere Asr
- S01. Qwen3-ASR-0.6B official card. URL: https://huggingface.co/Qwen/Qwen3-ASR-0.6B Inspected: 2026 release identity, Apache-2.0, language identification and 30-language/22-dialect coverage, separate forced aligner, inference wrapper, submodel versus whole-model size distinction. No vendor throughput result is adopted as a consumer-GPU measurement.
- S02. Qwen model metadata and immutable revision.
URL: https://huggingface.co/api/models/Qwen/Qwen3-ASR-0.6B
Inspected through authorized HTTPS fetch: revision
5eb144179a02acc5e5ba31e748d22b0cf3e303b0; safetensors metadata lists 938,008,576 BF16 parameters. Snapshot contains one model.safetensors plus config/processor/tokenizer files. Serialized count may differ from a deduplicated runtime parameter count; training prints the latter. Immutable card: https://huggingface.co/Qwen/Qwen3-ASR-0.6B/blob/5eb144179a02acc5e5ba31e748d22b0cf3e303b0/README.md
- S03. Qwen3-ASR technical report. URL: https://arxiv.org/abs/2601.21337 Inspected: official report abstract and model-family scope; no replication of the training corpus or benchmarks is claimed.
- S04. Cohere Transcribe 03-2026 official card. URL: https://huggingface.co/CohereLabs/cohere-transcribe-03-2026 Inspected: 2B, Apache-2.0, 14 languages, observed contact-sharing gate, native Transformers >=5.4.0 route, reported PyTorch 2.10.0 testing, missing timestamps/diarization, language and non-speech limitations. No access gate accepted. Weights not downloaded.
- S05. Cohere technical announcement. URL: https://cohere.com/blog/transcribe Inspected: Conformer encoder + Transformer decoder, waveform/log-Mel input, supervised token cross-entropy and from-scratch original training. These are model characteristics, not evidence that full training fits 24 GB.
- S06. Cohere Arabic official card and release note. URLs: https://huggingface.co/CohereLabs/cohere-transcribe-arabic-07-2026 ; https://docs.cohere.com/changelog/transcribe-arabic Inspected: official 2B Arabic/English fine-tune and Apache-2.0; introduction's code-switch optimization claim versus remaining code-switch caveat; observed access gate. No interpretation turns those caveats into a guarantee.
- S07. Cohere Embed and Rerank official current catalogs. URLs: https://docs.cohere.com/docs/cohere-embed ; https://docs.cohere.com/docs/rerank ; https://docs.cohere.com/docs/models Inspected: Embed v5.0 pro/fast and Rerank v4.0 pro/fast names and task roles. Public sources inspected here do not establish a downloadable training checkpoint for these exact products. No API account was created or called.
- S08. North Micro Vision Instruct official card. URL: https://huggingface.co/CohereLabs/North-Micro-Vision-Instruct Inspected: 2.4B total = 2B language + 400M vision; Apache-2.0; model-specific Transformers 5.16.0 support and linked training recipes. Used only to place the model in the right category and avoid treating a 2B language-backbone label as a total count.
- S09. Tiny Aya official card. URL: https://huggingface.co/CohereLabs/tiny-aya-global Inspected: explicit model-summary count 3.35B and CC-BY-NC-4.0. The card contains inconsistent boilerplate elsewhere; we rely on the model summary, not the erroneous 111B sentence in its terms section. No license permission beyond the card is inferred.
- S10. Command A+ official card. URL: https://huggingface.co/CohereLabs/command-a-plus-05-2026-bf16 Inspected: 218B total/25B active sparse MoE and Apache-2.0. Active parameters are not a storage or optimizer-memory estimate. Large-family conceptual comparison only.
- S11. Conformer foundational paper. URL: https://arxiv.org/abs/2005.08100 Inspected: convolution-augmented attention architecture concept. Historical reference for the mechanism, not an outdated checkpoint recipe.
- S12. Native Cohere ASR source at Transformers 5.4.0. URL: https://raw.githubusercontent.com/huggingface/transformers/v5.4.0/src/transformers/models/cohere_asr/modeling_cohere_asr.py Inspected: conditional generation forward accepts labels, shifts decoder inputs, computes supervised loss, and absorbs audio_chunk_index for generation. Trainability is source-supported; no fine-tune recipe or measured memory is claimed.
- S13. Official Cohere fine-tuning repository. URL: https://github.com/cohere-ai/cohere-finetune Inspected: supported base-model list and text-model LoRA/QLoRA scope. It does not establish Transcribe support. Mutable main inspected, not installed.
- S14. SciPy audio resampling. URL: https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.resample_poly.html Inspected: filtering and rational resampling semantics. Actual authoring tests ran SciPy 1.17.0; candidate training requirements pin 1.15.3. The project uses stable resample_poly and scipy.io.wavfile APIs, but the candidate environment was not installed.
- S15. Qwen model implementation at pinned repository revision. URL: https://github.com/QwenLM/Qwen3-ASR/blob/7c6daf77a2421100f5fb066495372c00129d39ff/qwen_asr/core/transformers_backend/modeling_qwen3_asr.py Inspected: thinker.audio_tower; audio embeddings replace placeholders; thinker.forward accepts labels/use_cache; causal loss; SDPA and gradient-checkpointing support; outer class primarily exposes generation. Original project calls thinker directly rather than copying the upstream monkey-patch.
- S16. Qwen processor implementation at pinned revision. URL: https://github.com/QwenLM/Qwen3-ASR/blob/7c6daf77a2421100f5fb066495372c00129d39ff/qwen_asr/core/transformers_backend/processing_qwen3_asr.py Inspected: chat template, audio-placeholder expansion, input_features/feature_attention_mask, default left padding. The project's one-example microbatch and prefix assertion deliberately avoid padded-prefix ambiguity.
- S17. CTC paper. URL: https://www.cs.toronto.edu/~graves/icml_2006.pdf Inspected: original connectionist temporal classification construction, blank labels and sum over valid alignments. Historical objective contrast only; the chapter does not claim its Qwen or Cohere loop uses CTC.
- S18. Transformers 4.57.6 causal loss implementation. URL: https://raw.githubusercontent.com/huggingface/transformers/v4.57.6/src/transformers/loss/loss_utils.py Inspected: causal label shifting, ignored -100 labels and mean loss. Project accumulates losses weighted by supervised shifted target count, rather than averaging unequal-length examples equally.
- S19. Official Qwen ASR fine-tuning guide/source. URLs: https://github.com/QwenLM/Qwen3-ASR/blob/7c6daf77a2421100f5fb066495372c00129d39ff/finetuning/README.md ; https://github.com/QwenLM/Qwen3-ASR/blob/7c6daf77a2421100f5fb066495372c00129d39ff/finetuning/qwen3_asr_sft.py Inspected: audio/text JSONL, language header, processor-driven supervised training and checkpoint support. This book provides an original bounded loop, not a verbatim copy or a claim that the upstream defaults fit 24 GB.
- S20. Official Qwen inference wrapper and audio utilities. URLs: https://github.com/QwenLM/Qwen3-ASR/blob/7c6daf77a2421100f5fb066495372c00129d39ff/qwen_asr/inference/qwen3_asr.py ; https://github.com/QwenLM/Qwen3-ASR/blob/7c6daf77a2421100f5fb066495372c00129d39ff/qwen_asr/inference/utils.py Inspected: from_pretrained kwargs; local paths; CPU device fallback; transcribe(language=..., return_time_stamps=False); 16 kHz normalization; optional separate forced aligner. CPU/GPU examples remain source-reviewed, not runtime-tested.
- S21. Qwen package metadata and repository commit.
URLs: https://github.com/QwenLM/Qwen3-ASR/blob/7c6daf77a2421100f5fb066495372c00129d39ff/pyproject.toml
; https://api.github.com/repos/QwenLM/Qwen3-ASR/commits/main
Inspected via public HTTPS: commit
7c6daf77a2421100f5fb066495372c00129d39ff, dated 2026-06-26; package version 0.0.6; exact Transformers 4.57.6, Accelerate 1.12.0 dependencies. Unpinned transitive dependencies remain the reason requirements-qwen.txt is labeled a candidate API target, not a solved lockfile.
- S22. PyTorch 2.10 automatic mixed precision. URL: https://docs.pytorch.org/docs/2.10/amp.html Inspected: autocast context and dtype behavior. The original loop explicitly uses FP32 model/optimizer storage and BF16 CUDA autocast. No GradScaler is used for that BF16 path; no claim of FP16 compatibility is made.
- S23. Transformers 5.4.0 Cohere documentation. URL: https://huggingface.co/docs/transformers/v5.4.0/model_doc/cohere_asr Inspected: native model/processor API and generation usage. Cohere local inference script syntax-checked only. Weight access is left to an authorized user-controlled setup.
- S24. Official PyTorch wheel installation matrix. URL: https://pytorch.org/get-started/previous-versions/ Inspected: PyTorch 2.10.0 Linux/Windows CUDA 12.6/12.8/13.0 and CPU wheel indexes. The chapter selects only torch for this project; torchvision/torchaudio are not required by its code. No package installation executed.
Large Architecture
- A01. DeepSeek-AI. Official DeepSeek-V3 inference
implementation, especially
MLA,Gate,ExpertandMoE. Inspected source; file revisionb15f0dbbbe6a4bc403306175698439ef380f5fb5, dated 27 August 2025. The optimizedabsorbpath storeskv_cacheandpe_cache; the code also exposes anaivepath. Pinned implementation
- A02. XiaomiMiMo. MiMo-V2.6-Flash-RL model
implementation. Inspected at checkpoint revision
5711b268169967567844e1e560e8a3966da959b1. Relevant classes:MiMoV2MoEGate,MiMoV2MoE,MiMoV2Attention,MiMoV2DecoderLayer, vision patch/merger and audio classes. The router'snoaux_tcbranch explicitly raises whenself.trainingis true. This establishes a limit of this implementation, not a claim that no other training system exists. Pinned source
- A03. Ainslie et al. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints, 2023. Primary definition and motivation for grouped-query sharing. The chapter does not reproduce the paper's speed/quality measurements as local results. Paper
- A04. DeepSeek-AI. DeepSeek-V3 Technical Report, first submitted 27 December 2024. Primary evidence for 671B total/37B active, MLA, MoE, 14.8T tokens and MTP. Historical reference, not a latest-release assertion. Paper
- A05. Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, 2022. Supports distinction between exact IO-aware execution and changing the attention graph. Paper
- A06. Dao and Gu. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, 2024. Primary Mamba-2/SSD paper. Paper
- A07. State Spaces. Official Mamba-2 module.
Inspected constructor, sequence path, step/cache handling and
convolution/state structure. File revision
6b72c122713bb769cc82c6b8e6d019c53d27d6a1, dated 7 October 2024. No kernel was executed. Pinned module
- A08. Kimi Team. Kimi Linear: An Expressive, Efficient Attention Architecture, 2025. Primary KDA reference; a gated recurrent-memory explanation, not a promise to reproduce the reported throughput. Paper
- A09. Xie et al. mHC: Manifold-Constrained Hyper-Connections, version 2, 5 January 2026. Supports constrained mixing and Sinkhorn normalization. The two-scalar worked example in the book is original, not an experimental result from the paper. Paper
- A10. XiaomiMiMo. Official MiMo-V2.6-Flash-RL card and file listing. The card reports 309B total/15B active, while the dynamic Hub tensor summary reports approximately 311B. Card license label is MIT. The listing contains configuration, implementation, technical-report PDF, modality components and weights. Existence and metadata were verified; weight contents were not downloaded or independently counted. Card and pinned files
- A11. LLM-Core Xiaomi. MiMo-V2.6: Scaling
Reinforcement Learning Towards Self-Improvement, technical-report PDF
supplied by the official checkpoint repository. Downloaded as a document
and text-extracted; architecture sections 2.1-2.4,
pretraining/mid-training section 3, router-freezing discussion and
post-training overview inspected. Report rounds Flash to 310B and Pro to
1.02T. It reports Flash 48T pretraining tokens (26T text stage +22T
multimodal stage) and Pro 30T (27T+3T). The chapter's discussion of the
report stays at the architecture/data-stage level. Pinned
technical report. Retrieved PDF SHA256:
fb81e6e083801b3358f084ed6be953dc23b0d2e434690f4541d5eae03e01e7af
- A12. XiaomiMiMo. Official MiMo-V2.6-Pro-RL card and
configuration. Checkpoint revision
73875d00b30a89ef8cc353a0b60b0e9f9561952d. Publisher architecture and rounded counts; code/config observations confirm layer pattern and head/expert geometry. Card, pinned configuration
- A13. Z.ai. Official GLM-5.3-Flash card. Publisher evidence for newly trained multimodal base, 320B total/18B active, sparse/linear hybrid, mHC and 30T corpus. Its MIT card label and linked paper are recorded without assuming publication of all training data. The official family repository separately lists FP8 and BF16 variants. Card, family repository
- A14. XiaomiMiMo. MiMo-V2.6-Flash-RL configuration,
revision
5711b268169967567844e1e560e8a3966da959b1. Directly counted 39 local and 9 global layers inhybrid_layer_pattern; inspected dimensions, expert settings, modality configs, max positions and quantization fields.quant_methodisfp8andstore_dtypeismxfp4; ignored tensors and mixed-precision components preclude an all-one-dtype byte claim. Pinned configuration
- A15. XiaomiMiMo. Separate DFlash draft configuration and implementation in the Flash-RL release. Five layers, window 1,024 and block size eight are in the separate config; it is not identified solely by the main model's MTP field. Pinned draft configuration, draft source
- A16. XiaomiMiMo. MiMo-V2.6-Flash-MOPD card.
Verified distinct post-training release and the card's teacher/prefix
and repetition-mitigation description. Checkpoint revision
2479e2d0029eca9a34cc7e7f55a121925f81908e; retrieved metadata created 27 September 2026. Its Flash config was inspected alongside RL config. The card also links Pro-MOPD. Card, pinned config
- A17. GLM-5 Team. GLM-5: from Vibe Coding to Agentic Engineering, arXiv 2602.15763. Verified as the report linked by the current official Flash card. Used only as a family reference, not as sole proof of the later Flash architecture. Paper
- A18. Z.ai. GLM-5.3-Flash configuration, revision
eb9eb208eb0d988989d07a6a12d0fdeb5f52574a. Direct observations: 34 linear/11 sparse-attention layers, 3 dense/42 MoE FFNs, 288 routed experts/top8+1 shared, four mHC streams, KDA geometry, index settings and modality config.indexer_typesall readfullin this release, even though supporting code allows reuse. Pinned configuration
- A19. Hugging Face Transformers.
modeling_glm5_next.py. Inspected actual KDA sequence/decode paths, FP32 recurrent-state storage, top-k router, mHC mixing, sparse indexer, compressed KV cache and expansion/reference attention path. File revision0a896aa41bba78f92338db21cc468fe043888657, dated 2 October 2026 at 03:23:05 UTC, before this research. It is a current source inspection, not a tested installed package. Pinned implementation
- A20. Hugging Face Transformers. GLM-5.3-Flash model
documentation. Direct statement that this implementation does not
include an MTP layer. The checkpoint's saved
transformers_versiondoes not guarantee that any arbitrary installed release supports every component. Documentation source
- A21. Qwen. Official Qwen3.5-397B-A17B card. 397B/17B, 60-layer 3-to-1 hybrid layout, MoE geometry and native-versus-extended context. Used as a selected architecture comparison rather than a latest-Qwen claim. Card
- A22. Qwen. Qwen3.5-397B-A17B configuration. Model
repository revision returned by metadata:
8472618112abcbd45acbcdc58436aff4233c23f7. Config read directly;layer_typesconfirms 45 linear and 15 full-attention layers. Pinned configuration
- A23. NVIDIA. cuSPARSELt data types and supported sparsity formats. Primary source for pattern-specific sparse-kernel contracts. This is not a performance claim for an unspecified consumer RTX. Documentation
- A24. Rajbhandari et al. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models, 2020 revision. Primary reference for reducing redundancy by sharding training states. The book does not extrapolate the paper's measured scaling to the user's hardware. Paper
- A25. PyTorch. FSDP/fully-shard documentation and NVIDIA Megatron Core parallelism documentation. Framework references for the distinction between model replicas and sharded parameters, tensor/pipeline/expert placement. PyTorch FSDP, Megatron Core. Verify runtime-specific meanings and support before implementing distributed work; no cluster recipe is promised in the book.
Training Lessons
- T01. micrograd's scalar engine. Andrej Karpathy, micrograd/engine.py.
Inspected the full engine, including graph traversal, backward closures,
repeated-use accumulation, and ReLU. Suitable before tensors become a
distraction. Educational scalar operations are not a performance model
for real training. Revision:
7bc720e951fe422b8f8814aa5aa1b64121d26b4c, default-branch commit 2026-08-03. Repository license: MIT. CPU learning exercise; no consumer-GPU benchmark claimed.
- T02. micrograd's correctness tests. test/test_engine.py, same revision. Inspected both tests comparing outputs and derivatives against PyTorch, including expressions that reuse values. Teaches an independent reference check and numeric tolerance. Tests were read, not run here; PyTorch was not installed for this review.
- T03. micrograd video and corrected companion.
Karpathy, original
YouTube lesson, published 2022-08-16; and second-half
lecture notebook. Access: author's YouTube description and chapter
list read on the original YouTube page; substantive companion notebook
code read. The notebook's
expbackward method contains an explicit correction from assignment to accumulation. Useful author chapter markers: 08:08, 1:22:28, 2:01:12. Course repository:karpathy/nn-zero-to-hero, MIT; revision73c3fcc741f0ec104ca850b1fb0df90e7e8d4cde, 2024-02-20. The course's correct repository name isnn-zero-to-hero, notneural-networks-zero-to-hero. Official syllabus states programming and introductory mathematics prerequisites; this book must teach that missing on-ramp itself.
- T04. an official visual explanation of gradient descent. Grant Sanderson, Gradient descent, how neural networks learn, lesson dated 2017-10-16; official text adaptation by Josh Pullen. Read the prediction-function/cost-function distinction, local gradient direction, learning-rate explanation, and held-out generalization discussion. Use for intuition after one concrete update. No figures or extended prose from it are reproduced in this book; the official written adaptation was read, not video playback.
- T05. makemore's MLP lecture notebook. Karpathy, makemore_part2_mlp.ipynb, course revision above. Read all code cells: whole-item train/dev/test partitioning, context construction, embedding lookup, minibatching, cross-entropy, manual updates, and sampling. Original lecture is linked by the creator's course README; its companion code, rather than captions, supplied the substantive evidence.
- T06. makemore as an executable educational trainer.
makemore
README. Read data format, scope, model choices, usage, default tiny
transformer, and source of the example names. It is a separate
executable project from the evolving lecture notebooks. Revision
988aa59e4d8fefa526d06f3b453ad116258398d4, default-branch commit 2022-11-20; MIT. The README is evidence for intended educational/CPU use, not an independently measured runtime.
- T07. diagnosing activation and gradient scale. makemore_part3_bn.ipynb, course revision above. Inspected training code, running statistics, gradient retention, histogram generation, saturation fraction, and update-to-weight statistics. Useful after a learner has a training loop to debug. The exercise's BatchNorm and tanh examples should not be generalized into a universal transformer architecture recipe.
- T08. corrected small GPT source for the video. ng-video-lecture/gpt.py.
Full file inspected. Its attention divides by the square root of the
key/head dimension and masks future positions. It demonstrates a clear
pedagogical decoder but is not a production serving or complete
checkpointing system. Revision
52201428ed7b46804849dea0b3ccf0de9df1a5c3, 2023-02-07. GitHub returned no license metadata for this repository; no reuse permission is inferred from its public visibility. This book links to it for study and uses independently written code. The final generation call is outside a no-grad/evaluation wrapper, so the book should retain its own explicit inference-mode hygiene rather than copying the script wholesale.
- T09. GPT video chapter markers and author corrections. Karpathy, original GPT lesson, published 2023-01-17. Read the original description, exercises, chapter list, and corrections on the original YouTube page. No claim of watching the entire video or reading a caption transcript. Useful markers: 14:27 for batches; 47:11 for weighted aggregation; 1:02:00 for learned self-attention; 1:26:48 for residual connections. Author corrections concern future-token masking at 57:00 and head-dimension scaling at 1:20:05. Code and paper are the authority for the corrected operations.
- T10. attention notation and its visual interpretation. Sanderson, Attention in transformers, step-by-step. Read its query/key/value explanation, notation warning, masking, dimensions, and illustrative-behavior caveat. Its column-oriented diagrams differ from the original paper's row-oriented convention. This is an excellent opportunity to teach shape annotations; attention visualizations should not be mistaken for a complete causal explanation of the trained model.
- T11. connect the implementation to the original Transformer. Vaswani et al., Attention Is All You Need, HTML v7. Read sections 3.2.1-3.2.3 and 5.2 for scaled attention, projections, masking, and the eight-P100 training premise. Suggested just-in-time question: which operation in the reader's code corresponds to each part of the attention equation? This is a focused paper-reading bridge; the architecture chapter also discusses this paper.
- T12. inspect a complete small training process. nanoGPT/train.py,
revision
3adf61e154c3fe3fca428ad6bc3818b27a3b8291, 2025-11-12; MIT. Inspected accumulation, mixed-precision scaling/clipping order, evaluation, checkpoint dictionary, and logging. Its readable structure is useful for process literacy. Its checkpoint fields do not establish complete deterministic replay; its final-microbatch log is not an exact accumulated-batch average.
- T13. deprecated examples are historical evidence. nanoGPT README, same revision. Read the deprecation notice, toy-run/reproduction distinction, and original hardware premise. Its GPT-2 reproduction used eight A100 40 GB GPUs for about four days. Include only as a brief caution about transferring old hardware/API assumptions, not as a recommended current trainer or model survey. The small character-model configuration was inspected to confirm it was a distinct experiment.
- T14. Raschka's minimal training loop and license
boundary. Sebastian Raschka, gpt_train.py,
requirements,
and license.
Read the standalone loop, evaluation mode, no-grad handling, token
counter, split, small batch/context settings, and dependencies. Revision
faa6205602f487c0382295c698ffafe3d635af39, 2026-10-01. The LICENSE.txt is based on Apache 2.0 but modifies the definition of source to exclude books and related images; GitHub metadata returnedNOASSERTION. Do not label the entire book/artwork corpus standard Apache-2.0 or reproduce it under the code's terms. Dependency lower bounds are not an exact environment lock.
- T15. make instruction-loss labels inspectable. gpt_instruction_finetuning.py,
same revision. Inspected dataset rendering and
custom_collate_fn, especially shifted targets, padding masks, retained terminal token, and truncation. The collator does not mask all instruction positions by default. The transferable lesson is to inspect the actual objective before making an “assistant-only training” claim.
- T16. LoRA as a loading and export workflow. Hugging
Face, smol-course
LoRA/PEFT lesson. Read the adapter-loading, merge, and trainer
examples. Revision
f445ae5d9dd83f355ce6a2b1c6c71b588c8e1541, 2026-09-17; Apache-2.0. A useful bridge after a learner understands which parameters receive gradients. The original LoRA paper was checked as a reading pointer; the LLM chapter supplies the detailed paper discussion.
- T17. an instructive tutorial-version mismatch.
smol-course SFT
lesson, hands-on
chapter, and root
requirements. Inspected actual configuration, formatting, optional
adapter, trainer, and upload blocks. The root lock names Torch 2.5.1,
Transformers 4.46.3, and TRL 0.12.1; the lesson uses newer model/API
examples. The ordinary trainer example's
args=configdiffers from the nearby definedtraining_config. Its full-fine-tuning snippets are not verified 24 GB recipes.
- T18. compare against the declared trainer version.
TRL
0.12.1 SFT documentation. Read its
max_seq_lengthconfiguration and completion-only collator discussion. This provides concrete evidence for checking API generation rather than combining contemporary examples with an old requirements file. Its token-context and EOS/padding cautions also show why a collator needs example-level inspection. Use the book's selected trainer version consistently.
- T19. embedding training as an evaluated retrieval
process. Sentence Transformers, NLI
example and MS
MARCO training example. Full scripts read. Revision
4a3b5cd6ec718e421f57e824a41ed3fd99595df6, 2026-09-21; Apache-2.0. Focus on initial evaluation, negative construction/filtering, duplicate avoidance, cached contrastive minibatches, and post-training evaluation. The MS MARCO script materializes corpus/query/teacher-score dictionaries in host RAM; caching does not eliminate those costs. Both inspected scripts attempt Hub upload despite “optional” comments. Their newer main-branch import layout should not be pasted into the embedding chapter's pinned v5.1.1 environment. Sentence-BERT is the theory bridge already covered in that chapter's source ledger.
- T20. toy denoising versus actual diffusion.
Jonathan Whitaker and Hugging Face, Diffusion
Models from Scratch notebook, creator's
companion lesson, and embedded original
walkthrough. Notebook Markdown/code cells read, including the
minimal UNet, uniform corruption, image-prediction objective, and
comparison with Gaussian DDPM, noise prediction, timestep conditioning,
and sampling. Video link verified as the companion page's embedded
lesson; no video/transcript viewing claimed and no timestamp invented.
Course revision
57b371aea6d477644726653c3318c3f36afe461c, 2026-09-17; Apache-2.0. The latest commit changes workflow maintenance, not necessarily the 2022-23 teaching content. Unpinned notebook installation cells are not a reproducible lock. The DDPM paper is already part of the diffusion bibliography.
- T21. follow audio representation through a training
loop. Diffusion
for Audio notebook, same course revision. Read code/Markdown for
resampling, random slicing, spectrogram conversion, batch construction,
denoising objective, reconstruction, and upload. Its
loss.backward(loss)call is a real source-code pitfall; use ordinary scalar backward unless an upstream-gradient weighting is intentional. Treat its pretrained model and music data as separately licensed inputs, not covered automatically by the course license. The generated model-card template'slicense: mitis not evidence that all resulting weights can legally receive that label.
- T22. check the mathematical meaning against
framework code. PyTorch 2.8, Tensor.backward,
confirms the optional argument is an upstream gradient and gradients
accumulate in leaves. Diffusers v0.40.0, DDPM
scheduler source, inspected
add_noiseand prediction-type branches. It uses square-root signal/noise coefficients and distinguishes epsilon, sample, and velocity prediction. These correct possible confusion between variance and standard-deviation notation in informal tutorials. The Audio Diffusion API confirms that this pipeline uses Mel-spectrogram representation and exposes reconstruction settings. It is a main-branch reference, not the book's environment lock.
- T23. llm.c teaches verification before low-level
optimization. Karpathy and contributors, llm.c
README and LayerNorm
Python reference. Inspected CPU/single-GPU distinctions, reference
testing, and the full LayerNorm forward/backward study. Revision
f1e2ace651495b74ae22d45d1723443fd00ecd3a, 2025-05-10; MIT. Useful after tensor training is understood. The CPU starter fine-tunes an existing GPT-2 checkpoint for 40 short-context steps; it is not a scratch-pretraining achievement.
- T24. nanochat gives an end-to-end process, with
larger hardware defaults. nanochat
README and CPU/MPS
script. README and complete small-run script inspected. Revision
92d63d4e8bb4df75c3b71618f31ddde2378b2bcd, 2026-07-03; MIT according to README. End-to-end tokenizer/pretrain/SFT/evaluation/inference structure is a useful advanced map. The speedrun targets eight H100 GPUs; the README requires adjustment below 80 GB. Author-reported price and runtime are not current retail quotes or RTX measurements.
- T25. current checkpoint state and dependency separation. nanochat checkpoint manager, base training script, and pyproject.toml, revision above. Read checkpoint serialization/loading, resume and data-loader-state plumbing, loop metadata, tokenizer compatibility check, and dependency declarations. Model and optimizer artifacts are separated; optimizer state is rank-specific for distributed runs. The inspected project pins Torch 2.9.1 and selects CPU or CUDA 12.8 wheels through mutually exclusive extras. This is its own environment, not a reason to mix its dependencies into the book’s other projects. No deterministic cross-device resume or RTX fit claim is inferred.