4 Project turn four sensor readings into a useful prediction
You have a row of readings: temperature, vibration, load, and age. You want either an alert about a future fault or an estimate of remaining operating time. This is a good first complete machine-learning project because the input and output are small enough to inspect. You can run it on a CPU, understand every transformation, and keep the same workflow when the inputs later become photographs or text.
The accompanying example creates fictional machines. Its numbers are educational, not an engineering reliability model. No external dataset, account, GPU, or paid service is needed.
Open a fresh terminal at the unpacked companion root. Use Python 3.12 for the tested CPU reference environment. Create and activate a dedicated environment before the first command that imports these packages. The activation command below is Bash; on Windows PowerShell use the corresponding .venv-cpu/Scripts/Activate.ps1 path.
cd examples/tiny-ml
python -m venv .venv-cpu
source .venv-cpu/bin/activate
python -m pip install -r requirements-cpu.txt
python train_tabular.py --task classification --out runs/classification
python predict_tabular.py runs/classification/model.joblib \
--values 72 3.5 0.9 36
The first command generates data, separates machine groups, trains three candidates, calibrates the chosen classifier, selects an alert threshold, evaluates a held-out test set, and saves a model. The second turns one new row into a probability and an alert. Read the output files before changing the model.
Verification status: the classification and regression programs were executed on CPU with Python 3.12.14, NumPy 2.3.5, scikit-learn 1.8.0, SciPy 1.17.0, and joblib 1.5.3. Exact metrics below describe that one synthetic run. Neural examples later in these chapters were syntax-checked, not trained or GPU-tested in the authoring environment.
Name the input the target and the decision
An input is the information available when a prediction must be made. Here it is a four-number vector, in this order:
- temperature in degrees Celsius
- vibration in millimeters per second
- load as a fraction between zero and one
- age in months
A target is what the model learns to predict. For classification it is a zero or one indicating whether the fictional fault occurred. A label is the recorded target attached to an individual example. The alert is a downstream decision: issue it when the predicted probability passes a threshold. Target, score, and action are different objects. A model can rank machines well yet use a poor alert threshold.
For real data, make the target definition more precise: “fault within the next seven days, excluding planned maintenance, using measurements collected before noon today.” Specify what happens to a machine whose next seven days have not yet been observed. Calling it a negative would silently introduce incorrect labels. Specify whether maintenance changes the outcome you are trying to predict.
The dataset contains five observations per machine. Machine ID is useful for separating groups, but is not an input feature. If we split individual rows randomly, readings from the same machine could appear on both sides of the evaluation. A model might benefit from machine-specific regularities that will not be available for a new machine.
The first design artifact is a prediction contract: input fields and units, prediction time, target definition, acceptable errors, and what happens when an input is missing. Do this before choosing a network.
How the architecture changes the data you need
A linear model assumes the useful relationship can be expressed with weighted features, perhaps after carefully justified feature transformations. A tree ensemble can learn local thresholds and interactions, but needs examples in the relevant regions of feature space. A deep tabular network adds capacity without supplying missing evidence. Start with the simpler candidates and plot validation performance as the number of independent training machines grows. A learning curve still improving with more machines suggests a data opportunity; a flat curve can indicate missing information, label noise, a weak representation, or an architecture limit. No universal row count guarantees success.
For this example, 2,000 rows are only 400 independent groups. For your task, count the independent units and the rare outcomes in each deployment-relevant subgroup. The architecture cannot learn the behavior of a regime absent from training just because its parameter count is large.
Start with a baseline you can beat
The example compares three systems:
- A constant classifier predicts the training-set fault rate for every row. It establishes what can be achieved without using the readings
- Logistic regression learns one coefficient per processed feature and converts the weighted sum into a probability. It is a useful compact linear baseline
- Histogram gradient boosting builds a sequence of small decision trees. Each new tree helps reduce the remaining prediction error. It can capture nonlinear thresholds and interactions without a neural network
In this dataset, the generator makes high load interact with risk through a threshold. A small tree ensemble has a reasonable opportunity to improve on a linear model. That is a hypothesis to test, not a promise that trees win every tabular problem. The scikit-learn classifier and regressor APIs describe the controls used in this example. M03 M04
The important input-to-output chain is:
raw row → missing-value treatment → numeric representation → model score
→ probability calibration → chosen threshold → alert
Logistic regression needs missing values filled in; the example uses the median measured on training rows and adds missingness flags. Standardization then subtracts the training mean and divides by the training standard deviation. Those operations are inside a pipeline, so the same fitted transformations travel with the model. The boosting implementation can handle missing values directly. Computing imputation values from the full dataset would let future evaluation data influence the model. M01
Separate fitting calibration selection and testing
At the default --machines 400, five readings per machine
produce 2,000 rows. The script shuffles machine IDs once and assigns
whole machines to four disjoint sets:
| Set | Rows in this run | Purpose |
|---|---|---|
| Training | 1,200 | Fit coefficients, trees, and preprocessing |
| Calibration | 300 | Fit a probability correction for the selected classifier |
| Validation | 200 | Choose the candidate and the alert threshold |
| Test | 300 | Report final behavior after the other choices are fixed |
These proportions are a teaching choice. They are not a universal prescription. On very small real datasets, grouped cross-validation can make better use of scarce examples, but all preprocessing and selection must then happen inside the appropriate training folds. When the product predicts future events, chronological evaluation may matter more than random group assignment. If the intended deployment is on new factories, holding out machines from the same factories is still insufficient: hold out factories too. M20
The script does not refit on the test set. It also disables the
boosting estimator's automatic internal early stopping; otherwise an
unexamined random row split inside the estimator could defeat the
carefully chosen group boundaries. With
early_stopping=False, max_iter is the
deliberately fixed tree-building budget.
Read the first result before tuning
The CPU run selected boosting. Its validation average precision was approximately 0.703, compared with 0.681 for logistic regression and 0.240 for the constant baseline. On the held-out test set it achieved:
- average precision: 0.693
- ROC-AUC: 0.788
- precision and recall at the selected threshold: both 0.633
- threshold: approximately 0.31
- confusion matrix: 177 true negatives, 33 false positives, 33 false negatives, 57 true positives
This is not “69.3% accuracy.” Average precision summarizes precision as recall changes while sweeping score thresholds. Precision asks what fraction of alerts were correct. Recall asks what fraction of actual faults were caught. ROC-AUC measures ranking between positive and negative cases; it can look comfortable while the alert workload is unacceptable on a rare-event problem. Read the confusion matrix in counts and decide whether the operational tradeoff is useful.
The threshold maximizes validation F1 in this exercise. F1 combines precision and recall; it is a convenient neutral classroom choice, not the correct business objective by default. If missing a fault costs much more than an unnecessary inspection, choose a validation threshold that meets a recall requirement or minimizes explicitly stated costs. If inspection capacity is ten machines per day, evaluate that policy directly. Never choose a threshold by trying values on the test labels.
A probability is a claim that needs checking
A prediction of 0.8 should mean that roughly eight out of ten
comparable cases assigned that probability are positive.
Calibration tests this claim. The example fits a
sigmoid calibrator on a held-out set using
CalibratedClassifierCV(FrozenEstimator(model), method="sigmoid").
It writes mean predicted probabilities and observed positive fractions
for reliability bins. Calibration data must be separate from the data
used to fit the original estimator. M02
Brier score and log loss also appear in the report, but neither isolates calibration by itself. Inspect bin counts and a reliability plot rather than interpreting one aggregate number as proof. A small calibration set produces noisy bins. Changes in fault prevalence, sensors, or maintenance policy can make yesterday's calibrated probabilities wrong tomorrow.
High confidence is not proof that an input is familiar. A model can be confidently wrong on out-of-distribution readings. Start with practical safeguards: reject malformed units, flag ranges absent from training, expose an “insufficient information” state, and review uncertainty-sensitive cases. An ensemble's disagreement can be informative, but is not a universal detector of novelty. Probability calibration does not transform an unreliable model into a reliable one.
What each tabular control changes
The command-line controls are intentionally short:
| Control | Default | What it means and how to change it |
|---|---|---|
--task |
classification |
Switch to regression for a continuous target; this
changes models, losses, and metrics |
--seed |
42 | Reproduces data generation and split assignment; examine several seeds after the basic pipeline works |
--machines |
400 | Number of independent machine groups; there are five rows per group |
--out |
runs/tabular |
Artifact destination; use distinct directories for experiments |
Inside the code, logistic C=1 is inverse regularization
strength: smaller values penalize large coefficients more.
max_iter=1000 is a solver iteration limit, not a desired
number of meaningful training epochs. Ridge regression's
alpha=1 is a coefficient penalty whose direction is the
opposite of logistic C: larger means stronger
regularization.
Boosting uses learning_rate=0.06, which controls each
tree's contribution; max_iter=120, the number of boosting
iterations; max_leaf_nodes=7, the maximum leaves in an
individual tree; min_samples_leaf=25, which discourages
tiny local rules; and l2_regularization=1, which penalizes
leaf values. Smaller learning rates often need more iterations. Deeper
or more numerous trees can fit more patterns and more noise. Do not
sweep every setting at once: inspect errors, form one hypothesis, change
one or two controls, and compare on validation. M03 M04
Artifacts are model.joblib, metrics.json,
split_machine_ids.json, and
test_predictions.csv. The model bundle includes
preprocessing where needed, feature order, task, and threshold. A
reload-and-predict equivalence check is part of the executed script.
Pickle-derived formats such as joblib can execute code when loaded: load
your own trusted artifact, not a random downloaded file. Record package
versions with the model. M24
Inference on the CPU and what the GPU changes
predict_tabular.py loads the whole trusted joblib
bundle, uses the saved preprocessing, preserves feature order, and
applies the saved threshold. Missing measurements can be supplied as
nan; do not replace them with zero unless zero has that
meaning. This scikit-learn workflow executes on CPU. There is no
meaningful “move it to CUDA” switch for these fitted estimators. A
GPU-specific replacement would be a different implementation requiring
compatibility and accuracy validation.
The example unseen row produces a fault probability about 0.885 and an alert in the executed CPU run. That prediction is about the fictional generator, not the physical world. Serving must validate units and feature ranges before calling the model.
Change the output predict a number instead
Run:
python train_tabular.py --task regression --out runs/regression
python predict_tabular.py runs/regression/model.joblib \
--values 72 3.5 0.9 36
The representation of the four inputs stays the same. The target becomes a fictional number of remaining hours. The constant baseline now predicts a mean; ridge regression predicts a regularized linear combination; the tree model predicts a continuous value. There is no class threshold or classifier probability calibration. The example reserves the calibration split for consistency but does not use it for this regression task.
Mean absolute error, or MAE, measures the average absolute discrepancy in the target's units. Root mean squared error, or RMSE, penalizes large mistakes more strongly. R-squared compares squared error against predicting the test-set mean; it can be negative. In the executed run, test MAE was 15.77 synthetic hours, RMSE was 20.10, and R-squared was 0.871. These numbers say nothing about actual machinery.
Changing target scale changes optimization. If one output is in millimeters and another in dollars, an unweighted sum of squared errors can let the larger numeric scale dominate. Standardize each target using training statistics, or choose weights that reflect actual costs; invert the transformation before reporting physical units. If you need a range rather than a point, train quantile predictors or use an appropriately validated prediction-interval method. A single mean estimate cannot express multiple plausible outcomes by itself.
Generalize the project to your own input to output task
The purpose of this lab is a reusable procedure, not a particular classifier. Write down one real input and the exact desired output. Then ask:
- Is the needed information present? A photograph of a sealed box does not determine its unseen contents. A noisy sensor snapshot may not determine the exact date a part will fail. More training cannot recover information that is absent. Change the input, predict a distribution, narrow the claim, or allow abstention
- Is a deterministic program better? Converting Celsius to Fahrenheit, checking a checksum, sorting records, and applying a published tax formula are rule-based computations. Implement and test the rule. Training a predictor adds avoidable approximation error
- What is an independent example? A customer, original image, patient, device, document family, or time period may generate many correlated rows. Split at the level matching deployment before augmentation or paraphrasing
- How will the target be represented? Choose the representation before choosing the loss. Preserve units, coordinate frames, missing-label masks, and any required ordering
- What failure actually matters? A correct class with the wrong boundary, an accurate average with catastrophic outliers, and a semantically similar but wrong policy page are different errors
- What must be measured outside the model? Latency, memory, privacy, human review load, input failures, and recovery behavior belong in the acceptance test
Here are useful mappings to extend the same process:
| Desired output | Representation | Starting objective | Important trap |
|---|---|---|---|
| One of K exclusive classes | K logits and one integer label | multiclass cross entropy | Softmax assumes the labels are mutually exclusive |
| Any subset of K labels | K independent logits and K binary labels | binary cross entropy per label | Missing annotations are not automatically negatives |
| One or several numbers | one or several continuous outputs | MAE, MSE, or Huber | Units and target scales change the objective |
| A bounded fraction | sigmoid output | a loss matching what the fraction represents | A proportion, event probability, and arbitrary bounded score have different statistical meanings |
| A pixel label map | class logits at every pixel | per-pixel cross entropy or binary BCE | Background imbalance can hide total foreground failure |
| Keypoint locations | coordinate vector or spatial heatmaps | coordinate regression or heatmap loss | Resize and crop transforms must update coordinates |
| A sequence of labels | one label distribution per input position | masked per-position cross entropy | Padding must not count as a real target |
| A variable-length sequence | token sequence with end marker | masked next-token objective | More than one output may be valid; exact match alone may mislead |
| An unordered set of objects | object slots with classes and geometry | matching plus class/geometry losses | Arbitrary label order creates contradictory supervision |
| A ranked list | query/document scores | pairwise or contrastive ranking loss | Unlabeled documents may still be relevant |
For a short sensor time window, your tensor could have shape
[batch, channels, time]. A one-dimensional convolution
learns local temporal patterns; an RNN or small attention model can use
longer dependencies. For a photograph plus numeric readings, encode the
image and standardized numbers separately, concatenate their
representations, and train a small output head. Compare each modality
alone against the combined system: a supposedly useful input may
contribute nothing or create a shortcut.
For a sequence output, use a padding mask and compute loss only on
valid positions. For multiple outputs, write
total_loss = task1_loss + weight * task2_loss and explain
the weight's units and intent. For an unordered output, use a model and
training objective that handle matching; do not force a network to learn
your arbitrary enumeration of objects.
A neural network becomes attractive when useful features are difficult to hand-design, when spatial or sequential structure matters, or when a suitable pretrained representation exists. On ordinary modest tabular datasets, keep the linear and tree baselines. “Use a GPU” is not the problem definition.
Checkpoint what you should now be able to do
You should be able to describe a new task without naming a model: the prediction-time inputs, target, independent split unit, baseline, error measure, and failure policy. If you cannot, stop before increasing model size.