Training Your
Own Models
Browse the book
Chapter 5 / 408 min read

5 Build a dataset that teaches the job

Start with twenty examples you can defend

Before collecting thousands of records, write twenty real examples of the task. For each, record the input available at prediction time, the desired output, why that output is correct, the source, and a grouping identifier. Read them as if you were a new annotator. If the target is unclear to a careful person, training is unlikely to resolve the ambiguity reliably.

Consider a support-message router with labels billing, delivery, and other. 'Where is my order?' is delivery. 'I was charged twice' is billing. 'Please cancel the order because it is late' exposes an ambiguity: should the router prioritize cancellation, delivery, or the team that owns the action? You need an annotation policy. Adding more examples without deciding the policy creates contradictory supervision.

Write the policy next to the data. Define each label, inclusion and exclusion cases, tie-breaking rules, treatment of missing information, and when to use an unknown or abstain label. Include examples that distinguish easily confused categories. A label definition that only restates its name is not enough.

For a generated response, the target policy may specify facts that must be included, claims that must not be invented, formatting, length, tone, and what to do when evidence is absent. For an image mask, it must specify which pixels belong to the object and how to treat occlusion. For audio, it must specify segment boundaries, sample rate, loudness handling, and whether silence is meaningful.

A practical record format

JSON Lines stores one JSON object per line. It is convenient because a program can read records one at a time, and one malformed line can be identified precisely. Here is a task-neutral pattern:

{"id":"ticket_0042","group_id":"customer_17","input":"I was charged twice","target":"billing","source":"owned_support_export","split":"train"}

The id identifies the record. group_id identifies related records that must be handled together when splitting. input is what the model receives. target is what it should learn to produce. source records provenance. split indicates the assigned partition. Keep extra metadata outside the model-visible input unless it will genuinely be available at prediction time.

The eventual trainer may require a different schema, such as messages for chat fine-tuning or image and text fields for diffusion. Maintain a clean canonical dataset and a reproducible conversion step. Do not repeatedly edit the only copy to satisfy each new trainer's expectations.

A schema validator should check required fields, types, allowed labels, empty values, unique IDs, valid paths, and length limits. It should fail loudly on malformed records rather than silently skipping a large fraction. Log rejected records with a reason, taking care not to expose private contents unnecessarily.

Split related things before making derivatives

Training data teaches parameters. Validation data guides experiment choices. Test data supports a final estimate after choices are fixed. The exact percentages depend on dataset size and task diversity; a mechanical 80/10/10 split is not always appropriate. A test set of ten rare cases may be too small to tell you much, regardless of its percentage.

The split unit is often larger than one row. Keep messages from the same conversation together. Keep crops and augmented versions of the same image together. Keep clips from the same recording session together. Keep questions derived from the same document together when the intended test is generalization to new documents. Otherwise nearly identical information can appear on both sides of the evaluation boundary.

A time-based split is useful when the deployed system will face future data. A source-based split tests transfer to new sources. A person- or device-based split can test whether the system works beyond the individuals or equipment seen during training. Pick the split that matches the intended claim.

Deduplication must operate across the eventual boundaries. Exact hashes catch identical bytes, but near duplicates may differ only in whitespace, a timestamp, a crop, or a paraphrase. Use normalization appropriate to the modality and inspect suspiciously similar pairs. Do not normalize away meaningful distinctions, such as case in identifiers or punctuation in code, merely to increase duplicate counts.

Data leakage also occurs when preprocessing learns from evaluation data. A scaler's mean, a feature selector's decisions, a learned tokenizer for a strict from-scratch experiment, or a vocabulary fitted to all records may incorporate information from outside the training partition. Fit learned preprocessing on training data, then apply it unchanged to validation and test data. F09

Teach the difficult cases on purpose

A dataset should represent both common situations and important failure modes. If 95 percent of support messages are delivery questions, a model that always predicts delivery can look accurate while being useless for billing. Count examples per class and per important slice. A slice might be language, source, message length, image lighting, audio noise level, or tool type.

Do not correct imbalance mechanically before asking what deployment will look like. Oversampling a rare class changes how often its examples influence training. Class weights change how much its errors contribute to loss. Thresholds change decision behavior after prediction. These are different interventions. Evaluate under the actual operating distribution as well as on deliberately balanced diagnostic sets.

Negative examples matter. A retrieval model needs plausible irrelevant passages, not only obviously unrelated ones. A tool-calling assistant needs cases where no tool is appropriate and cases where required arguments are missing. A visual classifier needs backgrounds and confusing non-target objects. A segmentation model needs empty scenes if it will encounter them. Without such examples, the model may learn that every input demands the positive behavior.

Hard negatives are examples that resemble a correct case but are wrong for a specific reason. They can be powerful, but they must be correctly labeled. A supposed negative that is actually relevant damages the learning signal. Review a sample manually before scaling mined negatives.

Synthetic data is a tool with a contract

Synthetic data can isolate a concept, cover a rare structured case, or create cheap preliminary tests. The first classifier in this book uses it because the exact rule is known. The image and audio exercises use original procedural examples so you can learn the pipeline without scraping uncertain material.

Synthetic data is not automatically representative. A classifier may learn the generator's punctuation or naming conventions rather than the intended meaning. A model trained on generated answers can inherit the generator's errors and stylistic habits. Ten paraphrases of one source are not ten independent sources.

When using a larger model to generate training material, preserve the prompt, model/version, sampling settings, source evidence, validation procedure, and applicable terms. Verify factual and structural claims rather than trusting fluency. Do not allow the generator to see held-out answers and then treat its outputs as independent training examples. Keep related generated variants in one split group.

Use a small human-reviewed anchor set to assess whether generated material teaches the intended behavior. Increase synthetic volume only after the first trained comparison improves on real held-out cases. If performance improves only on synthetic tests from the same generator, report that narrower result.

Rights privacy and consent are dataset properties

Record where each source came from, what rights or permission support its use, whether redistribution is allowed, and whether derivative model artifacts have additional conditions. Publicly accessible content is not automatically unrestricted training material. An open-weight model may have a license with conditions different from an open-source software license.

Separate the questions of permission to train, permission to publish data, permission to publish weights, and permission to deploy a commercial service. They can have different answers. Read the exact model and dataset terms for the pinned revision and intended use. This book provides a workflow for checking those terms, not a legal determination for a particular dataset.

Minimize personal information. Remove identifiers that the task does not need. Use de-identified or synthetic records for early experiments. A model may memorize unusual training strings, so removing a private source file after training does not establish that the information is absent from the weights. If strict deletion or access control is required, a retrieval store with explicit records may be easier to govern than parameter updates.

For voices and identifiable people, obtain appropriate consent for the intended use. Do not assume permission to record implies permission to train a generative model or imitate a person. For children or sensitive domains, additional safeguards and legal obligations may apply; use authoritative guidance for the relevant jurisdiction and project.

Build a data card before the long run

A useful data card states the task, source types, collection dates, rights, consent basis where applicable, preprocessing, split method, counts, length distributions, label distribution, known gaps, and intended uses. It should also state prohibited or unsupported uses and who reviewed the labels.

Save the exact processed files or reproducible transformation recipe. Compute hashes of split files. If you later change cleaning or labels, assign a new dataset version. 'The same dataset' is not a sufficient experiment description when its contents have changed.

Inspect at least a small random sample and a deliberately difficult sample after preprocessing, not only before it. Decoding errors, truncation, broken chat templates, swapped label IDs, and misaligned image masks often appear in the processed representation. Print or render examples at the point where the model will consume them.

Exercises

Create twenty examples for a task you care about. Find two that could reasonably receive different labels from different annotators. Rewrite the policy until the disagreement can be resolved consistently, or add an explicit uncertain category.

Choose a grouping rule. Explain why splitting individual rows would be misleading for your data. Then state exactly what your held-out result will claim: new messages from known users, new users, future documents, new recording sessions, or something else.

Write one positive example, one easy negative, one hard negative, and one case that should trigger abstention. If you cannot identify the abstention case, your deployment boundary is probably underspecified.