32 From small model to small device product
A 24 GB training GPU and a phone have different constraints. A model that trains comfortably on your workstation may exceed a phone's RAM, run slowly on its supported kernels, or heat the device under continuous use. In embedded practice, TinyML often means microcontroller-class deployment with very small memory and power budgets; the examples above are small models, not a claim that every artifact fits every microcontroller.
Measure the entire path on the target device
Write an acceptance budget before optimizing:
- model file size and installation/download size
- peak runtime memory, including inputs, activations, buffers, tokenizer, and retrieval index
- cold-start latency, steady-state median latency, and tail latency
- sustained throughput, thermal behavior, and battery cost
- supported operators and data types on the chosen CPU, GPU, or NPU backend
- input preprocessing, output decoding, and any privacy or offline requirements
Parameter count is only one term. The current project's 1024-dimensional FP32 embedding uses 4,096 bytes before index overhead: one million such document vectors require about 4.096 decimal GB. A tiny encoder plus a huge local index can be a poor phone design. Keep only the required corpus, compress and re-evaluate, or choose a permitted server-side architecture when appropriate.
Export quantize and verify instead of assuming
A typical deployment path is train → choose checkpoint → export with fixed preprocessing contract → optimize for the target backend → validate numerical parity → evaluate the optimized model → benchmark on actual hardware. ExecuTorch is PyTorch's current on-device ecosystem for mobile and edge workloads; follow its version-specific exporter and backend documentation rather than copying an old mobile tutorial blindly. This chapter does not claim a tested phone export. M19
Quantization represents some values with fewer bits. Post-training quantization uses a representative calibration set to choose ranges; quantization-aware training simulates the effect during training. Neither guarantees preserved accuracy. Calibration for numeric quantization ranges is different from probability calibration. Test thin segmentation structures, rare classes, long texts, and low-confidence cases after conversion, not just average agreement on easy examples.
For simple models, reducing image resolution, feature channels, sequence length, or index size may matter more than an exotic optimization. Each reduction changes what information remains available. Distillation can train a smaller student to imitate a stronger teacher, but the student still needs representative examples and independent evaluation. A successful export proves format compatibility, not product quality.
A compact failure diagnosis checklist
| Symptom | First checks | Next controlled experiment |
|---|---|---|
| Training loss never falls | inspect targets, dtypes, non-finite inputs, trainable parameters, gradient flow | overfit a handful of examples |
| Training improves; validation worsens | leakage-free split, label consistency, sample size, distribution shift | smaller model, stronger justified regularization, earlier checkpoint |
| All predictions are majority class/background | label counts, loss reduction, missing-label handling, output decoding | foreground/class-aware objective and honest per-class metrics |
| Good IoU, bad outline | resolution, coordinate mapping, thin features, boundary tolerance | boundary-sensitive evaluation and higher-resolution patches |
| Good retrieval loss, worse search | false negatives, duplicate positives, too-easy evaluation, index/model mismatch | inspect hard examples and compare the frozen encoder |
| GPU out of memory | longest input, activation sizes, batch, evaluation buffering, other processes | one representative step at a smaller batch/resolution/length |
| Export succeeds, output changes | normalization, channel order, padding, unsupported operator fallback, quantization | compare intermediate or final tensors on a saved test bundle |
| Offline scores good, field behavior poor | subgroup/session/time shift, input acquisition, unseen cases, actual latency | gather rights-cleared field examples and rebuild the evaluation contract |
The end-to-end skill is now the same across these projects: define what information is available, represent the target faithfully, start with a cheap baseline, fit only on authorized training data, choose decisions on validation, test once on meaningful held-out cases, and verify the exact deployed pipeline. The architecture follows from that contract.