Training Your
Own Models
Browse the book
Chapter 9 / 407 min read

9 Project locate one object with a box

What detection adds to classification

An image classifier answers what kind of image it sees. A detector also answers where a target is. Common application patterns include locating a part before measurement, proposing a crop for another model, finding an object in a fixed inspection station, and reporting an empty scene without inventing a detection. A downstream application still decides what to do with the box.

This project trains an original small CNN to detect at most one foreground object. It returns a probability that an object is present and one enclosing rectangle. It deliberately does not solve crowded multi-object scenes, identify individual overlapping instances, or generate a variable-length list of classes and boxes. Those require a more capable detector and appropriate annotations. This bounded project teaches the complete data-to-box path before you study those systems.

The model is small and trained from scratch. Its usefulness comes from the tightly scoped synthetic task, not from a pretrained historical architecture. You can replace the synthetic examples with a similarly bounded real inspection task only after reviewing whether the representation, labels, and input variation match.

Make boxes from masks without changing the meaning

Open a fresh terminal at the companion root, enter examples/tiny-ml, and activate the neural environment installed in the image-classification project. If you skipped that project, complete its explicit torch and torchvision installation block first.

cd examples/tiny-ml
source .venv-neural/bin/activate
python make_shapes.py --out data/one-object --size 128 \
  --train 256 --val 64 --test 64 --seed 51 \
  --max-objects 1 --empty-probability 0.20
python detector_common.py

The generator writes an RGB image and binary mask for each example. A binary mask is a same-sized image whose background pixels are 0 and foreground pixels are 255; it records where the target exists. About twenty percent are empty by the generator's sampling rule; the exact split counts vary. Every nonempty image contains at most one generated object. The training, validation, and test sets use separate seeds. Their simple distribution is a teaching fixture, not evidence of robustness to real photographs.

A positive mask becomes a bounding box. Find the minimum and maximum foreground x and y coordinates. The maximum edge is one pixel beyond the largest included coordinate, making it exclusive. Divide x coordinates by image width and y coordinates by image height. A box [0.2, 0.2, 0.7, 0.6] therefore spans those fractions of the original image dimensions.

An empty mask has objectness target zero. Its placeholder box is ignored by the localization loss. Teaching the model that every image contains an object would produce a poor empty-scene behavior. The generator's max-objects setting is saved in the manifest, and the trainer refuses data not explicitly generated for this one-object contract.

If you derive one box from a mask containing several disconnected objects, you obtain one box around their union. That is a region-extent task, not ordinary instance detection. Do not make that substitution silently. The supplied detector uses max-objects 1 precisely to keep the target meaning clear.

Why the network retains a spatial grid

The input is an RGB tensor [3, 96, 96] after bilinear resizing and scaling pixel values to [0,1]. Three convolution/ReLU/pooling stages produce 16, 32, and 64 feature channels. A 4 by 4 pooled grid is retained and flattened before the output layer. Unlike a global average over the entire image, this retains coarse information about where features occur.

Direct layer-shape arithmetic gives 28,709 trainable parameters; the program records the authoritative runtime count. This is intentionally tiny, so the first challenge is correct data and output semantics rather than filling GPU memory.

The head emits five values. One is an objectness logit. Four are transformed to values between zero and one and arranged as ordered lower and upper x/y corners. Ordering ensures the decoded rectangle has nonnegative width and height. It does not guarantee that the rectangle is correct.

This architecture needs varied object locations, sizes, backgrounds, and empty scenes. If every target were centered, the network could learn a constant box rather than localization. If every positive had one background color and every negative another, it could detect the background. The synthetic generator varies nuisance conditions while preserving a learnable foreground cue.

For a real task, collect separate acquisition sessions and keep images of the same physical object together when splitting. Include the sizes, viewpoints, and absences that deployment will encounter. A one-object architecture is appropriate only when the application's field of view genuinely enforces that constraint.

Two losses teach two outputs

Objectness uses binary cross-entropy on the raw logit. Localization uses smooth L1 loss between predicted and target normalized box coordinates, only on positive examples. The total is objectness loss plus five times localization loss. Five is an explicit balancing choice for this teaching problem, not a universal detection constant.

This differs from classification's one class label per image and segmentation's one class label per pixel. The desired output representation determines what supervision is needed and where the loss applies. A box label is cheaper than a pixel mask for many real tasks, but it does not teach the precise outline inside the rectangle.

The optimizer is AdamW with learning rate 0.001, weight decay 0.0001, and the usual beta defaults of 0.9 and 0.999. Gradient norm is clipped at 1.0. The reference runs in FP32 on both CPU and CUDA. The seed is 42. These choices make the small experiment inspectable; tune on validation only if you replace its data.

Train save and evaluate

python train_detector.py --data data/one-object \
  --out runs/one-object-cpu --epochs 2 --batch-size 8 \
  --lr 0.001 --device cpu

# After the smoke test, a separate GPU experiment
python train_detector.py --data data/one-object \
  --out runs/one-object-gpu --epochs 15 --batch-size 16 \
  --lr 0.001 --device cuda

The CPU run is a short correctness pilot. A compatible CUDA setup can run the same computation on the GPU. The trainer refuses unavailable CUDA rather than silently falling back. It refuses a nonempty output directory so an old artifact cannot be overwritten by a different experiment.

--epochs controls complete passes over training images. --batch-size controls concurrent images per update. --lr controls update scale. --data locates the immutable split directories and manifest. --out identifies a fresh run. --device chooses the execution target. Input size, feature widths, box-loss weight, class mapping, threshold, data hashes, and numeric policy are saved explicitly.

The checkpoint with the lowest validation objective becomes best.pt. config.json contains the architecture identity, foreground class name, preprocessing, box convention, and threshold. metrics.json contains the epoch history and final held-out test results. These files form one inference artifact and must stay together. This small trainer intentionally does not implement optimizer resume; use a new run for a new experiment.

Evaluation reports mean IoU over all positive images, detection precision and recall requiring IoU at least 0.5, and false-alarm rate on empty images. A positive image with a high objectness score but the wrong box is an incorrect detection. The IoU threshold and objectness threshold are different: one measures geometric agreement; the other decides whether to emit a box. The default objectness threshold is a teaching choice of 0.5, not a calibrated production threshold.

Load on CPU and GPU with a fresh image

Generate a separate small set that is not used for fitting or model selection:

python make_shapes.py --out data/one-object-fresh --size 128 \
  --train 1 --val 1 --test 1 --seed 999 \
  --max-objects 1 --empty-probability 0

python predict_detector.py runs/one-object-gpu \
  data/one-object-fresh/test/images/00000.png \
  --device cpu --overlay outputs/detection-cpu.png

python predict_detector.py runs/one-object-gpu \
  data/one-object-fresh/test/images/00000.png \
  --device cuda --overlay outputs/detection-gpu.png

The inference script loads config.json and the trusted tensor state, reconstructs the network, applies the saved RGB/96-pixel preprocessing, enters evaluation mode, and moves the model and input to the selected device. It returns objectness probability and, when the threshold is met, the class name and box in original-image pixel coordinates. The optional overlay draws that box on the original image.

Both CPU and GPU use FP32 here. Small numerical differences are possible, so compare probabilities, boxes, and threshold crossings before claiming equivalence. A tiny single-image workload may not run faster on the GPU once transfer and launch overhead are included. Measure your actual latency and batching pattern.

Check what succeeded and what remains untested

During book preparation, the one-object dataset generator was executed, mask-to-box and empty-mask conversions were tested, and numerical IoU/detection-metric assertions passed. All new Python files compiled. PyTorch training and inference were not executed in the authoring environment, so no detection quality, GPU speed, or memory measurement is claimed.

Before a larger run, inspect several images and their derived boxes. Overfit a handful of clean examples, save and reload, then test an empty image and a fresh positive. If the model always emits a central box, inspect positional variation and the retained spatial grid. If it finds objects but returns poor boxes, inspect coordinate order, normalization, image resizing, and localization-loss weight.

The next segmentation project predicts a detailed mask rather than a rectangle. Keep that distinction concrete: a good enclosing box can contain much background, and a good class prediction says nothing about boundaries.