Training Your
Own Models
Browse the book
Chapter 29 / 406 min read

29 Scale a working experiment toward one billion parameters

The next experiment should answer one question

After the tiny transformer works, the tempting next step is to make it a thousand times larger. A better next step asks a specific question: is the small model underpowered for a task whose data and evaluation are already trustworthy? If you cannot show that the task needs more capacity, increasing size mainly increases the cost of uncertainty.

A useful scaling experiment keeps data processing, split definitions, and evaluation fixed while changing one capacity dimension. Compare a slightly wider model, a deeper model, or a longer training budget. Measure quality, memory, and throughput. The result should tell you whether another increase is justified.

Before scaling, replace the toy corpus with sufficient lawful material for the intended task. Repeating 133,000 synthetic bytes millions of times does not create the linguistic coverage of a real pretraining corpus. You may obtain a model that memorizes the generator extremely well and generalizes poorly.

Three different meanings of full training

Training from scratch initializes all learned parameters without a pretrained checkpoint and optimizes them on your data. Full-parameter fine-tuning starts from pretrained weights and allows all selected model parameters to change. Continued pretraining starts from pretrained weights and continues a broad self-supervised objective, often on domain text. These can all update every parameter, but their starting capability, data needs, and goals differ.

If you fine-tune a 600-million-parameter pretrained model, you inherit the results of its original large-scale pretraining. A 600-million-parameter model trained from scratch on your small dataset does not inherit that capability. Comparing their parameter counts without their data and training histories is misleading.

Head-only training and adapter training update a subset of parameters. They are valuable techniques, but should not be described as full-parameter training. A model can have a billion total parameters and only a few million trainable parameters. Record both counts.

A realistic ladder of model sizes

The educational byte transformer begins below one million parameters. Models in the low millions are suitable for learning mechanics, constrained synthetic tasks, and some genuinely narrow applications. Tens of millions allow more capacity while keeping many experiments relatively manageable. Hundreds of millions become a meaningful systems and data project. Near one billion, full training on 24 GB requires careful numeric, activation, optimizer, and sequence-length choices even when persistent state appears to fit.

Do not read those categories as performance rankings. A compact classifier with the right features may outperform a much larger generative model on a fixed decision task. A pretrained 350-million-parameter encoder or decoder may outperform a randomly initialized billion-parameter network because its representation already contains useful structure.

The companion architecture is intentionally easy to count. With vocabulary V, context T, width d, and L blocks, its tied-embedding parameter count is approximately Vd + Td + L(12d squared + 9d) + 2d. The exact implementation determines the bias terms. The program prints the authoritative count from distinct parameter tensors.

For illustration, width 768, twelve blocks, a 256-byte vocabulary, and context 512 give roughly 86 million parameters. Width 1,792 with twenty-four blocks and the same small vocabulary gives roughly 927 million. These are shape calculations, not recommended pretrained architectures or measured 24 GB training configurations. A real subword vocabulary, different feed-forward ratio, untied output head, or grouped-query attention changes the count.

Do not scale all costs at once

Increasing width changes many matrices quadratically. Increasing depth adds repeated blocks. Increasing vocabulary enlarges embedding and output layers. Increasing context raises activation and attention costs. Increasing batch raises concurrent activation cost. If you change all of them together, you lose the ability to diagnose the bottleneck.

For a larger full-training feasibility test, start with microbatch one and a short but meaningful sequence. Run complete updates, including optimizer-state allocation. Then measure longer sequences and evaluation. Only after the full lifecycle works should you consider accumulation to reach a chosen effective batch.

A short sequence can make a model fit while making the task impossible. If the input evidence and target depend on a 2,000-token document, a 128-token feasibility demonstration does not establish that the useful task fits. Resource constraints and task requirements must be reconciled honestly.

One billion is a project boundary rather than a magic number

Under the simple FP32 AdamW inventory, one billion trainable parameters require about 16 GB of persistent state before activations. A 24 GB card may have enough remaining capacity for some short-context configurations, but not all. Mixed precision, optimizer implementation, attention kernels, output-logit materialization, and checkpointing determine the real peak.

An efficient optimizer or CPU offload can move the boundary. It also changes numerical behavior, throughput, CPU RAM demand, and failure modes. A memory-saving flag should be accompanied by an explanation of which state it reduces and a new measured smoke test. If the implementation keeps an FP32 master copy, include it in the inventory.

The practical full-training projects in this book stay at or below one billion parameters. The supplied from-scratch model is much smaller so the reader can understand and validate it. The pretrained full-fine-tuning project uses a small published checkpoint to demonstrate a useful adaptation path without pretending to reproduce its original pretraining.

What the three billion ceiling actually means

Three billion parameters is within the model-selection discussion, including adapter methods and specialized architectures. It is not a blanket promise of ordinary full-parameter training in 24 GB. In the simple FP32 AdamW case, the persistent training state alone is about 48 GB. Lower-precision and offloaded methods must be assessed individually.

A three-billion-parameter four-bit base may fit for adapter training because the frozen base has compressed storage and only a small subset needs optimizer state. That does not make three-billion-parameter full training equivalent. Likewise, a mixture-of-experts name mentioning three billion active parameters may have far more resident parameters.

Pretraining from scratch adds a compute and data constraint even if memory can be solved. A model that sees too little varied text may never develop useful general language ability. A long run on inadequate data can be less useful than a much shorter adaptation of a well-chosen pretrained model.

Build a measured budget before committing

Measure useful tokens per second over a representative interval after startup. Include a separate end-to-end figure accounting for evaluation and checkpointing. Estimate total token presentations from data size, epochs or sampling plan, and masking. Divide by measured throughput, then add contingency for restarts and preprocessing.

Write down the quality decision that will justify continuing. For example: after the first fixed budget, the larger model must beat the smaller model on held-out task accuracy without exceeding a latency limit. If it does not, inspect errors before spending more. Scaling can help with capacity; it cannot repair a mislabeled task or a missing input field.

A sensible project budget also includes human work: reviewing data, defining labels, inspecting failures, and maintaining the resulting artifact. GPU time is only one part of the cost. For many personal projects, a day spent improving data and evaluation is more valuable than a day spent training a larger network.

A scaling experiment worksheet

Record the hypothesis, baseline configuration, proposed change, exact parameter count, trainable parameter count, numeric policy, estimated persistent state, measured peak state, sequence or image/audio shape, effective batch, data version, token or sample budget, measured throughput, estimated duration, stopping criteria, and evaluation result.

Keep the first larger run deliberately bounded to feasibility and early learning. Do not interpret a successfully completed ten-step run as evidence of eventual quality. It answers a systems question. Quality needs a separate training budget and held-out evaluation.

Exercises

Using the stated parameter formula, calculate how doubling width differs from doubling depth. Explain why width can make memory grow much faster than expected.

Write two project plans for the same domain: one from-scratch tiny model and one pretrained full fine-tune. List which capability each starts with, what data it needs, and what result would justify its cost.

A three-billion-parameter checkpoint loads successfully in four-bit inference. List the additional facts you need before claiming full-parameter training will fit: trainable dtypes, master copies, optimizer state, gradients, activations, context, batch, working buffers, and supported training kernels.