Training Your
Own Models
Browse the book
Chapter 34 / 404 min read

34 Design your first independent model

Choose a job small enough to verify

Your first independent project should have a clear input, a limited output, examples you can lawfully use, and a result you can inspect. 'Build my own general assistant' contains too many unresolved tasks. 'Route these five kinds of repair request and abstain when unsure' is a project. 'Answer questions from this manual with a cited passage' is a project. 'Generate short original machine sound effects in three categories' is a project.

Choose a task where you can obtain correct targets and where mistakes are recoverable. Avoid making a first experiment responsible for medical, legal, financial, hiring, or other high-impact decisions. Learning the machinery and establishing real-world validity are separate obligations.

Write the prediction contract

Describe the exact input fields or file format, what is available at prediction time, the output format, the target policy, and the downstream action. State which outputs are valid and what happens when required information is missing. List examples outside the intended domain.

Then identify the application pattern. A fixed decision suggests classification or a typed-decision model. A numeric quantity suggests regression. A ranked list suggests retrieval or ranking. A pixel mask suggests segmentation. Flexible text suggests conditional generation. Images or audio synthesized from conditions suggest an appropriate generative model. A deterministic transformation may need ordinary code rather than training.

Do not force a problem into the category whose tutorial you most recently read. The input/output contract is the organizing principle.

Build a baseline and a small evaluation set

Create a baseline with the simplest plausible method. Save its outputs on a carefully selected initial evaluation set. Include common cases, costly errors, missing information, and domain boundaries. Define a primary metric, guardrails, and a stopping rule.

The evaluation set should be large enough to expose important failures and small enough that you can review it. It can expand as the project matures. Preserve a final holdout that does not guide repeated choices. Record sources and grouping units from the beginning so later expansion does not create leakage.

If the baseline already meets the requirement, deployment engineering may be more valuable than training. A model project succeeds when the user's problem is solved, even if the winning method is simpler than expected.

Select the smallest useful intervention

Try a better representation or prompt before a major weight update when appropriate. Try retrieval for changing factual knowledge. Try a small head or adapter when the pretrained representation is useful but the task interface differs. Try full fine-tuning when you have a reason to update the whole model and can afford the controlled experiment. Train from scratch when the objective, architecture, data ownership, or narrow task makes that justified.

For every choice, write why the data should teach the architecture. If the model predicts next tokens, show how the desired behavior appears in target sequences. If it scores options, show positive choices and plausible alternatives. If it predicts masks, show aligned pixel labels. If it denoises audio, show consistent waveform preprocessing and conditioning.

Use four gates

The data gate requires a valid schema, clear targets, lawful provenance, representative coverage, and protected split boundaries. The correctness gate requires a tiny run, finite gradients, a parameter update, and save/load equivalence. The feasibility gate requires measured peak memory, throughput, storage, and recovery under representative shapes. The quality gate requires a held-out improvement over the baseline without unacceptable regressions.

Do not skip a gate because the previous one was exciting. A clean dataset does not prove a trainer is correct. A working trainer does not prove the useful workload fits. A feasible run does not prove the result is good.

Plan the next experiment from errors

After the first trained comparison, classify the failures. Missing evidence suggests input or retrieval changes. Ambiguous labels suggest annotation work. Format errors suggest target consistency or constrained decoding. Long-input failures suggest truncation or context design. Confident wrong decisions suggest calibration and abstention work. Systematic visual or audio artifacts suggest representation, preprocessing, or training-coverage problems.

Choose one change that addresses the dominant error category. Write the predicted effect and how you will measure it. Preserve the prior artifact so you can compare or roll back. This creates a series of understandable experiments rather than a pile of increasingly complicated configurations.

A final handoff to yourself

At the end, you should have an inference-ready artifact, a readable model card, a data card, a tested CPU or GPU deployment path, an evaluation report, an environment record, and a list of known limitations. You should also have one next decision: deploy within the stated boundary, improve a particular weakness, or stop because the approach is not worth its cost.

The central skill is no longer knowing one command. It is connecting a desired behavior to evidence, representation, architecture, optimization, and evaluation. That skill transfers from three learned numbers to hundreds of millions of parameters, and it lets you read larger-model research without confusing an impressive architecture with a feasible personal project.