Training Your
Own Models
Browse the book
Chapter 30 / 403 min read

30 Understand what a public training process really demonstrates

You can learn a great deal from a repository without being able to reproduce its biggest run. The useful habit is to inspect the relationship among code, configuration, data and reported result.

Audit a public fine tuning script before running it

Hugging Face’s inspected smollm/text/finetuning/train.py is a useful example. It exposes a model, dataset, sequence length, accumulation, LoRA settings and optional NF4 loading. But the inspected file still passes max_seq_length to SFTConfig, while reviewed TRL 1.14.1 uses max_length. It also defaults to Hub upload and Weights & Biases reporting. A reputable repository can contain a script written for an older API and defaults inappropriate for private data. L52 L06

The lesson is not to avoid public examples. Before running one, answer:

  1. Which commit and dependency versions does this script expect?
  2. What data does it load, and are you allowed to use it?
  3. Is the loss all-token, completion-only or assistant-only?
  4. What is actually trainable, and how is precision configured?
  5. Does it upload data, weights or logs by default?
  6. Which evaluation produced the reported result?
  7. Can you run one batch, save, reload and generate before the long job?

Our examples use local fixtures, explicit masks and disabled reporting/uploading. The implementation was written independently for this book rather than presented as an unmodified official benchmark recipe.

Compare scratch training with adapting a finished model

Training from scratch initializes weights randomly and must learn basic token relationships, language patterns and task-relevant structure. Full fine-tuning starts from pretrained weights but updates all of them. “Full” answers which weights change; “from scratch” answers where the weights came from. A full fine-tune is not equivalent to recreating the original model.

The inspected SmolLM3 model card reports 11 trillion pretraining tokens and 384 H100 GPUs. Its public pretraining README describes a 2.36-million-token global batch and 24 days on those 384 GPUs, along with training/configuration resources. This is a transparent published process worth studying. It is not a one-consumer-GPU recipe. L49 L53

For your hardware, scratch training a tiny transformer is an excellent learning project. Scaling an educational run to tens or hundreds of millions of parameters is possible only with a matching data, memory and time budget. A sub-billion parameter ceiling does not make general-purpose pretraining cheap. Your tiny model can demonstrate real learning without matching the broad competence of a publicly pretrained checkpoint.

A common dense-transformer planning approximation is training_FLOPs ≈ 6 N D, for N parameters and D training tokens. It omits important architecture, attention and recomputation details, so use it for order-of-magnitude reasoning. The compute-optimal training literature studies data/model allocation under particular budgets; a remembered “tokens per parameter” ratio is not a universal data requirement or a guarantee of useful quality. L54

For an explicitly hypothetical example, N=1e9 and D=20e9 gives 1.2e20 FLOPs. At an assumed sustained 50 TFLOP/s, division gives about 27.8 days of uninterrupted arithmetic. That throughput is not a measured RTX result and the token count is not a recommendation; data processing, checkpointing, inefficiency and experiments add time. The calculation simply shows why a model that fits in memory can still be expensive to pretrain.

For a real estimate, measure processed training tokens per second on your actual architecture and sequence length, then use time ≈ planned_tokens / measured_tokens_per_second, adding evaluation and saving overhead. Separate processed input tokens from supervised target tokens when reporting SFT throughput. An hour with short answers and heavily masked prompts is not directly comparable to an hour of all-token pretraining.

A stop rule that matches the experiment

A smoke test stops after establishing finite loss, valid masks, a complete update, saving and reloading. A useful supervised experiment stops when validation no longer improves under the chosen protocol or the pre-agreed token/compute budget is spent. A final held-out test estimates how well the selected configuration transfers.

Keep a simple result record: model and revision, data hashes and split method, trainable count, optimizer/precision assumptions, effective tokens per update, peak memory, throughput, validation trend, held-out task scores and a few representative failures. Report a measured improvement only after making that comparison. Honest small results teach more than an impressive model name attached to an unverified recipe.