7 Move from feature columns to neural representations
The sensor project used a short feature vector. Images and token sequences need larger structured representations. Before running the next image project, learn the two pieces that change: tensor shapes and learned layers.
Scalars vectors matrices and tensors
A scalar is one number. A vector is an ordered list of numbers. A matrix is a rectangular table of numbers. Tensor is the broader term used for arrays with any number of dimensions. In a training program, a tensor also has a numeric type and a device.
Shape tells you how its values are arranged. A batch of 32 records, each with 10 features, has shape [32, 10]. A batch of 8 RGB images resized to 224 by 224 pixels commonly has shape [8, 3, 224, 224] in PyTorch: batch, channel, height, width. A batch of 4 token sequences, each of length 256, has shape [4, 256]. After a language model converts token IDs to 128-dimensional vectors, the representation has shape [4, 256, 128].
The order of dimensions matters. A tensor can have the right number of values and the wrong meaning. Passing images arranged as [batch, height, width, channel] into code expecting [batch, channel, height, width] may fail loudly or produce nonsensical behavior. Record shapes near the boundary of each stage: load, preprocess, batch, forward, loss, and decode.
A tensor's dtype is its numeric format. Token IDs are integers. Weights and activations are usually floating-point values. A Boolean mask indicates which positions are valid or allowed. A device indicates where the tensor lives, such as CPU or cuda:0. A model and its inputs generally need compatible devices. Moving a model to a GPU does not automatically move every separately created input tensor there.
A neural layer is a learned transformation
A linear layer maps an input vector x to an output vector y using a matrix W and an optional bias b. Written compactly, y = Wx + b. If the input has 10 features and the output has 20 features, the layer has 200 weights plus 20 biases. It learns combinations of the input features.
Stacking linear layers without anything nonlinear between them is still equivalent to one linear transformation. A nonlinear activation, such as ReLU or GELU, allows a network to express more complicated relationships. ReLU keeps positive values and replaces negative values with zero. GELU makes a smoother change. You need not memorize their formulas to understand the role: they prevent a deep network from collapsing into a single linear map.
Depth is the number of successive transformations. Width is the size of intermediate representations. More depth or width often increases capacity, the variety of functions the network can express. Capacity is not a guarantee of useful learning. A larger network can memorize bad labels, amplify dataset shortcuts, or remain poorly trained because the available data and compute are insufficient.
A residual connection adds an earlier representation to the output of a transformation. It lets a block learn a modification rather than requiring it to reconstruct everything from scratch. Normalization layers rescale intermediate representations using a particular rule. Their details differ, but their practical role includes keeping numerical behavior manageable. When adapting a known architecture, preserve these details unless architecture research is your explicit project.