Training Your
Own Models
Browse the book
Chapter 4 / 4013 min read

4 Project turn four sensor readings into a useful prediction

You have a row of readings: temperature, vibration, load, and age. You want either an alert about a future fault or an estimate of remaining operating time. This is a good first complete machine-learning project because the input and output are small enough to inspect. You can run it on a CPU, understand every transformation, and keep the same workflow when the inputs later become photographs or text.

The accompanying example creates fictional machines. Its numbers are educational, not an engineering reliability model. No external dataset, account, GPU, or paid service is needed.

Open a fresh terminal at the unpacked companion root. Use Python 3.12 for the tested CPU reference environment. Create and activate a dedicated environment before the first command that imports these packages. The activation command below is Bash; on Windows PowerShell use the corresponding .venv-cpu/Scripts/Activate.ps1 path.

cd examples/tiny-ml
python -m venv .venv-cpu
source .venv-cpu/bin/activate
python -m pip install -r requirements-cpu.txt
python train_tabular.py --task classification --out runs/classification
python predict_tabular.py runs/classification/model.joblib \
  --values 72 3.5 0.9 36

The first command generates data, separates machine groups, trains three candidates, calibrates the chosen classifier, selects an alert threshold, evaluates a held-out test set, and saves a model. The second turns one new row into a probability and an alert. Read the output files before changing the model.

Verification status: the classification and regression programs were executed on CPU with Python 3.12.14, NumPy 2.3.5, scikit-learn 1.8.0, SciPy 1.17.0, and joblib 1.5.3. Exact metrics below describe that one synthetic run. Neural examples later in these chapters were syntax-checked, not trained or GPU-tested in the authoring environment.

Name the input the target and the decision

An input is the information available when a prediction must be made. Here it is a four-number vector, in this order:

  1. temperature in degrees Celsius
  2. vibration in millimeters per second
  3. load as a fraction between zero and one
  4. age in months

A target is what the model learns to predict. For classification it is a zero or one indicating whether the fictional fault occurred. A label is the recorded target attached to an individual example. The alert is a downstream decision: issue it when the predicted probability passes a threshold. Target, score, and action are different objects. A model can rank machines well yet use a poor alert threshold.

For real data, make the target definition more precise: “fault within the next seven days, excluding planned maintenance, using measurements collected before noon today.” Specify what happens to a machine whose next seven days have not yet been observed. Calling it a negative would silently introduce incorrect labels. Specify whether maintenance changes the outcome you are trying to predict.

The dataset contains five observations per machine. Machine ID is useful for separating groups, but is not an input feature. If we split individual rows randomly, readings from the same machine could appear on both sides of the evaluation. A model might benefit from machine-specific regularities that will not be available for a new machine.

The first design artifact is a prediction contract: input fields and units, prediction time, target definition, acceptable errors, and what happens when an input is missing. Do this before choosing a network.

How the architecture changes the data you need

A linear model assumes the useful relationship can be expressed with weighted features, perhaps after carefully justified feature transformations. A tree ensemble can learn local thresholds and interactions, but needs examples in the relevant regions of feature space. A deep tabular network adds capacity without supplying missing evidence. Start with the simpler candidates and plot validation performance as the number of independent training machines grows. A learning curve still improving with more machines suggests a data opportunity; a flat curve can indicate missing information, label noise, a weak representation, or an architecture limit. No universal row count guarantees success.

For this example, 2,000 rows are only 400 independent groups. For your task, count the independent units and the rare outcomes in each deployment-relevant subgroup. The architecture cannot learn the behavior of a regime absent from training just because its parameter count is large.

Start with a baseline you can beat

The example compares three systems:

  • A constant classifier predicts the training-set fault rate for every row. It establishes what can be achieved without using the readings
  • Logistic regression learns one coefficient per processed feature and converts the weighted sum into a probability. It is a useful compact linear baseline
  • Histogram gradient boosting builds a sequence of small decision trees. Each new tree helps reduce the remaining prediction error. It can capture nonlinear thresholds and interactions without a neural network

In this dataset, the generator makes high load interact with risk through a threshold. A small tree ensemble has a reasonable opportunity to improve on a linear model. That is a hypothesis to test, not a promise that trees win every tabular problem. The scikit-learn classifier and regressor APIs describe the controls used in this example. M03 M04

The important input-to-output chain is:

raw row → missing-value treatment → numeric representation → model score
        → probability calibration → chosen threshold → alert

Logistic regression needs missing values filled in; the example uses the median measured on training rows and adds missingness flags. Standardization then subtracts the training mean and divides by the training standard deviation. Those operations are inside a pipeline, so the same fitted transformations travel with the model. The boosting implementation can handle missing values directly. Computing imputation values from the full dataset would let future evaluation data influence the model. M01

Separate fitting calibration selection and testing

At the default --machines 400, five readings per machine produce 2,000 rows. The script shuffles machine IDs once and assigns whole machines to four disjoint sets:

Set Rows in this run Purpose
Training 1,200 Fit coefficients, trees, and preprocessing
Calibration 300 Fit a probability correction for the selected classifier
Validation 200 Choose the candidate and the alert threshold
Test 300 Report final behavior after the other choices are fixed

These proportions are a teaching choice. They are not a universal prescription. On very small real datasets, grouped cross-validation can make better use of scarce examples, but all preprocessing and selection must then happen inside the appropriate training folds. When the product predicts future events, chronological evaluation may matter more than random group assignment. If the intended deployment is on new factories, holding out machines from the same factories is still insufficient: hold out factories too. M20

The script does not refit on the test set. It also disables the boosting estimator's automatic internal early stopping; otherwise an unexamined random row split inside the estimator could defeat the carefully chosen group boundaries. With early_stopping=False, max_iter is the deliberately fixed tree-building budget.

Read the first result before tuning

The CPU run selected boosting. Its validation average precision was approximately 0.703, compared with 0.681 for logistic regression and 0.240 for the constant baseline. On the held-out test set it achieved:

  • average precision: 0.693
  • ROC-AUC: 0.788
  • precision and recall at the selected threshold: both 0.633
  • threshold: approximately 0.31
  • confusion matrix: 177 true negatives, 33 false positives, 33 false negatives, 57 true positives

This is not “69.3% accuracy.” Average precision summarizes precision as recall changes while sweeping score thresholds. Precision asks what fraction of alerts were correct. Recall asks what fraction of actual faults were caught. ROC-AUC measures ranking between positive and negative cases; it can look comfortable while the alert workload is unacceptable on a rare-event problem. Read the confusion matrix in counts and decide whether the operational tradeoff is useful.

The threshold maximizes validation F1 in this exercise. F1 combines precision and recall; it is a convenient neutral classroom choice, not the correct business objective by default. If missing a fault costs much more than an unnecessary inspection, choose a validation threshold that meets a recall requirement or minimizes explicitly stated costs. If inspection capacity is ten machines per day, evaluate that policy directly. Never choose a threshold by trying values on the test labels.

A probability is a claim that needs checking

A prediction of 0.8 should mean that roughly eight out of ten comparable cases assigned that probability are positive. Calibration tests this claim. The example fits a sigmoid calibrator on a held-out set using CalibratedClassifierCV(FrozenEstimator(model), method="sigmoid"). It writes mean predicted probabilities and observed positive fractions for reliability bins. Calibration data must be separate from the data used to fit the original estimator. M02

Brier score and log loss also appear in the report, but neither isolates calibration by itself. Inspect bin counts and a reliability plot rather than interpreting one aggregate number as proof. A small calibration set produces noisy bins. Changes in fault prevalence, sensors, or maintenance policy can make yesterday's calibrated probabilities wrong tomorrow.

High confidence is not proof that an input is familiar. A model can be confidently wrong on out-of-distribution readings. Start with practical safeguards: reject malformed units, flag ranges absent from training, expose an “insufficient information” state, and review uncertainty-sensitive cases. An ensemble's disagreement can be informative, but is not a universal detector of novelty. Probability calibration does not transform an unreliable model into a reliable one.

What each tabular control changes

The command-line controls are intentionally short:

Control Default What it means and how to change it
--task classification Switch to regression for a continuous target; this changes models, losses, and metrics
--seed 42 Reproduces data generation and split assignment; examine several seeds after the basic pipeline works
--machines 400 Number of independent machine groups; there are five rows per group
--out runs/tabular Artifact destination; use distinct directories for experiments

Inside the code, logistic C=1 is inverse regularization strength: smaller values penalize large coefficients more. max_iter=1000 is a solver iteration limit, not a desired number of meaningful training epochs. Ridge regression's alpha=1 is a coefficient penalty whose direction is the opposite of logistic C: larger means stronger regularization.

Boosting uses learning_rate=0.06, which controls each tree's contribution; max_iter=120, the number of boosting iterations; max_leaf_nodes=7, the maximum leaves in an individual tree; min_samples_leaf=25, which discourages tiny local rules; and l2_regularization=1, which penalizes leaf values. Smaller learning rates often need more iterations. Deeper or more numerous trees can fit more patterns and more noise. Do not sweep every setting at once: inspect errors, form one hypothesis, change one or two controls, and compare on validation. M03 M04

Artifacts are model.joblib, metrics.json, split_machine_ids.json, and test_predictions.csv. The model bundle includes preprocessing where needed, feature order, task, and threshold. A reload-and-predict equivalence check is part of the executed script. Pickle-derived formats such as joblib can execute code when loaded: load your own trusted artifact, not a random downloaded file. Record package versions with the model. M24

Inference on the CPU and what the GPU changes

predict_tabular.py loads the whole trusted joblib bundle, uses the saved preprocessing, preserves feature order, and applies the saved threshold. Missing measurements can be supplied as nan; do not replace them with zero unless zero has that meaning. This scikit-learn workflow executes on CPU. There is no meaningful “move it to CUDA” switch for these fitted estimators. A GPU-specific replacement would be a different implementation requiring compatibility and accuracy validation.

The example unseen row produces a fault probability about 0.885 and an alert in the executed CPU run. That prediction is about the fictional generator, not the physical world. Serving must validate units and feature ranges before calling the model.

Change the output predict a number instead

Run:

python train_tabular.py --task regression --out runs/regression
python predict_tabular.py runs/regression/model.joblib \
  --values 72 3.5 0.9 36

The representation of the four inputs stays the same. The target becomes a fictional number of remaining hours. The constant baseline now predicts a mean; ridge regression predicts a regularized linear combination; the tree model predicts a continuous value. There is no class threshold or classifier probability calibration. The example reserves the calibration split for consistency but does not use it for this regression task.

Mean absolute error, or MAE, measures the average absolute discrepancy in the target's units. Root mean squared error, or RMSE, penalizes large mistakes more strongly. R-squared compares squared error against predicting the test-set mean; it can be negative. In the executed run, test MAE was 15.77 synthetic hours, RMSE was 20.10, and R-squared was 0.871. These numbers say nothing about actual machinery.

Changing target scale changes optimization. If one output is in millimeters and another in dollars, an unweighted sum of squared errors can let the larger numeric scale dominate. Standardize each target using training statistics, or choose weights that reflect actual costs; invert the transformation before reporting physical units. If you need a range rather than a point, train quantile predictors or use an appropriately validated prediction-interval method. A single mean estimate cannot express multiple plausible outcomes by itself.

Generalize the project to your own input to output task

The purpose of this lab is a reusable procedure, not a particular classifier. Write down one real input and the exact desired output. Then ask:

  1. Is the needed information present? A photograph of a sealed box does not determine its unseen contents. A noisy sensor snapshot may not determine the exact date a part will fail. More training cannot recover information that is absent. Change the input, predict a distribution, narrow the claim, or allow abstention
  2. Is a deterministic program better? Converting Celsius to Fahrenheit, checking a checksum, sorting records, and applying a published tax formula are rule-based computations. Implement and test the rule. Training a predictor adds avoidable approximation error
  3. What is an independent example? A customer, original image, patient, device, document family, or time period may generate many correlated rows. Split at the level matching deployment before augmentation or paraphrasing
  4. How will the target be represented? Choose the representation before choosing the loss. Preserve units, coordinate frames, missing-label masks, and any required ordering
  5. What failure actually matters? A correct class with the wrong boundary, an accurate average with catastrophic outliers, and a semantically similar but wrong policy page are different errors
  6. What must be measured outside the model? Latency, memory, privacy, human review load, input failures, and recovery behavior belong in the acceptance test

Here are useful mappings to extend the same process:

Desired output Representation Starting objective Important trap
One of K exclusive classes K logits and one integer label multiclass cross entropy Softmax assumes the labels are mutually exclusive
Any subset of K labels K independent logits and K binary labels binary cross entropy per label Missing annotations are not automatically negatives
One or several numbers one or several continuous outputs MAE, MSE, or Huber Units and target scales change the objective
A bounded fraction sigmoid output a loss matching what the fraction represents A proportion, event probability, and arbitrary bounded score have different statistical meanings
A pixel label map class logits at every pixel per-pixel cross entropy or binary BCE Background imbalance can hide total foreground failure
Keypoint locations coordinate vector or spatial heatmaps coordinate regression or heatmap loss Resize and crop transforms must update coordinates
A sequence of labels one label distribution per input position masked per-position cross entropy Padding must not count as a real target
A variable-length sequence token sequence with end marker masked next-token objective More than one output may be valid; exact match alone may mislead
An unordered set of objects object slots with classes and geometry matching plus class/geometry losses Arbitrary label order creates contradictory supervision
A ranked list query/document scores pairwise or contrastive ranking loss Unlabeled documents may still be relevant

For a short sensor time window, your tensor could have shape [batch, channels, time]. A one-dimensional convolution learns local temporal patterns; an RNN or small attention model can use longer dependencies. For a photograph plus numeric readings, encode the image and standardized numbers separately, concatenate their representations, and train a small output head. Compare each modality alone against the combined system: a supposedly useful input may contribute nothing or create a shortcut.

For a sequence output, use a padding mask and compute loss only on valid positions. For multiple outputs, write total_loss = task1_loss + weight * task2_loss and explain the weight's units and intent. For an unordered output, use a model and training objective that handle matching; do not force a network to learn your arbitrary enumeration of objects.

A neural network becomes attractive when useful features are difficult to hand-design, when spatial or sequential structure matters, or when a suitable pretrained representation exists. On ordinary modest tabular datasets, keep the linear and tree baselines. “Use a GPU” is not the problem definition.

Checkpoint what you should now be able to do

You should be able to describe a new task without naming a model: the prediction-time inputs, target, independent split unit, baseline, error measure, and failure policy. If you cannot, stop before increasing model size.