Training Your
Own Models
Browse the book
Chapter 33 / 405 min read

33 Turn a checkpoint into a usable application

The model is one component of the system

A useful application validates input, applies the exact preprocessing used in training, runs the model, decodes output, checks constraints, and decides what happens next. A trained checkpoint alone does not perform all of those responsibilities. Deployment is the process of preserving that contract outside the training script.

For a sensor predictor, the contract includes feature names, order, units, missing-value policy, calibration, and alert threshold. For a language model, it includes tokenizer, chat template, context limit, stopping behavior, and generation settings. For image or audio generation, it includes the base model, adapter, conditioning, scheduler, preprocessing, and output conversion.

Keep deterministic checks outside the model when they can be explicit. A tool argument can be checked against a schema. A date can be parsed. A path can be restricted to an allowed directory. A financial or destructive action can require approval. Model confidence is not permission to bypass such boundaries.

Package the smallest complete inference artifact

For each project, create an export directory containing the required weights or adapter, architecture configuration, preprocessing or tokenizer files, class mapping or output schema, dependency specification, and a short model card. Include a small known input and expected output or tolerance-based test.

For an adapter, identify the exact base model and revision. If redistribution of the base is not allowed or not desired, document how the authorized user obtains it. Do not imply that an adapter-only folder is self-contained. A merged model may simplify inference but increases size and must respect the relevant licenses.

A classifier saved with joblib or pickle-like serialization should be loaded only from trusted sources because such formats can execute code. Tensor-focused formats can reduce some risks, but preprocessing code and remote model code still need review. Avoid enabling trust_remote_code casually; it permits repository-provided Python to run in your environment.

CPU and GPU are two deployment targets

CPU inference is useful for small classifiers, many embedding workloads, low-throughput services, and portability. GPU inference is useful when model size, latency, batch throughput, or generation workload justify it. The same artifact may support both, but the runtime and dtype choices can differ.

Do not assume a GPU-oriented four-bit training configuration has a CPU inference path. A quantization library may require specific hardware kernels. A model can sometimes be exported to a different CPU-supported format, but that is a new conversion with its own quality and compatibility tests. Loading full FP32 weights on CPU may be simpler for a small model, at the cost of RAM and latency.

CPU memory is still finite. One billion FP32 weights alone require about 4 GB before runtime overhead and caches. A larger context, multiple concurrent requests, and tokenizer or preprocessing buffers add more. GPU serving similarly needs headroom for caches and concurrency rather than just the model's weights.

Benchmark the actual request pattern. A batch of 100 classification records measures throughput; one interactive request measures latency. Report startup time separately from steady-state inference. Include preprocessing and decoding in an end-to-end measurement, and distinguish it from model-only time.

Test export equivalence

Choose a fixed small suite of representative inputs. Run the original selected checkpoint and the exported artifact with matching preprocessing and deterministic settings where appropriate. Compare class probabilities or logits within a justified tolerance, generated outputs under fixed decoding, or task-level metrics when small numerical differences are expected.

For a merged adapter, first compare the unmerged base-plus-adapter path with the merged full-precision path. Then evaluate any quantized conversion. This staged comparison identifies where a regression appears. Converting and quantizing simultaneously can make diagnosis harder.

A model that loads successfully has passed a structural test. It still needs semantic tests. A wrong class-name order can preserve every tensor shape while reversing the application's decisions. A changed chat template can make a language model behave differently without any loading error.

Keep the application bounded

Specify maximum input sizes and output sizes. Reject malformed data clearly. Decide how to handle empty text, corrupt images, unsupported audio rates, missing fields, and requests outside the trained domain. A silent fallback to an arbitrary default can be more dangerous than an explicit failure.

For tool use, separate proposal from execution. Let the model propose a structured action. Let deterministic code validate the tool name, argument types, allowed values, permission, and current state. Use a dry-run or sandbox in testing. High-impact actions need the appropriate human or policy approval independent of the model's score.

For retrieval-assisted generation, preserve source identifiers and permission checks. A model should not retrieve a document merely because an embedding search finds it relevant if the current user lacks access. Fine-tuning does not provide per-document access control.

Monitor without quietly changing the task

Track production input distributions, error reports, abstention rates, latency, and relevant outcome metrics. A drift in data does not automatically justify retraining. First determine whether the issue is a changed target, changed population, preprocessing bug, or missing coverage.

Collect feedback with clear provenance. User corrections may be useful labels, but they can also be mistaken, malicious, or inconsistent. Do not automatically train on every interaction. Review, deduplicate, protect privacy, and version the resulting dataset.

Keep a rollback path. A new model should be evaluated against the old one before replacement. Deploy gradually when the application warrants it, and retain enough information to reproduce a reported error. The ability to restore a known-good artifact is part of reliability.

Write a model card that makes honest claims

State the model's purpose, architecture and base revision, training method, trainable parameter count, data sources and rights, split methodology, evaluation results, hardware and software used, tested inference targets, known limitations, and unsupported uses. Distinguish measured results from expected behavior.

A personal model card can be short, but it should still answer what changed from the base and how you know. If GPU training was not run during preparation of a recipe, say so. If a result is based only on synthetic data, say so. If a model was tested on one language or one recording condition, do not imply broad coverage.

Exercises

Package one model so a fresh Python process can load it and predict a provided sample without access to the training notebook's variables. List every required file.

Measure ten cold starts and a sequence of warm predictions on the intended CPU or GPU. Report the conditions and distinguish model-only from end-to-end timing.

For a proposed tool action, write the deterministic validator before asking a model to produce the action. Explain which conditions can be enforced by code and which require human judgment.