2 Your computer and the training workspace
What the parts of the computer do
The CPU runs the operating system, Python, many preprocessing steps, and some parts of your data loader. System RAM holds ordinary program memory. Your SSD stores source data, processed data, model downloads, checkpoints, and outputs. The GPU performs large batches of numeric operations. Its VRAM is a separate pool of memory with different constraints from system RAM.
A computer with 64 GB of RAM and a 24 GB GPU does not have an 88 GB GPU. A program may explicitly move some state between CPU RAM and VRAM, a technique called offloading. Transfer takes time and is limited by the connection between the devices. Offloading may make a run possible while making it much slower. It must be configured by a library or your own code; it is not a transparent extension of VRAM.
The GPU may also drive your monitor, browser, and other applications. A card sold as 24 GB does not necessarily offer a training process every byte of that nominal capacity. Drivers and kernels consume memory. Other processes consume memory. Allocators reserve blocks. Transient operations can require a higher peak than the steady state visible between steps. Leave a measured margin rather than planning a job to use precisely the advertised capacity.
A model download is not the complete disk budget. You may need the original checkpoint, a local cache, processed shards, several training checkpoints, a merged model, a quantized export, generated evaluation outputs, and temporary conversion files at the same time. A small adapter can coexist with a much larger frozen base model. Plan storage by enumerating artifacts, not by looking at the adapter file alone.
A terminal is a way to give explicit instructions
A terminal runs a shell, which interprets commands. A command consists of a program name and arguments. In the command below, python is the program, train.py is a file to run, --steps names an option, and 20 supplies its value.
python train.py --steps 20
The command acts relative to a current working directory. If the script expects data/train.jsonl, it means the data directory beneath that working directory unless the code says otherwise. Use pwd on Linux or macOS to print the current directory and ls to list files. In PowerShell, Get-Location and Get-ChildItem serve similar purposes. Paths with spaces must be quoted. An absolute path begins at the filesystem root or a drive; a relative path begins at the current directory.
The examples in this book use Bash-style commands and forward-slash paths. Linux is a straightforward reference platform for the common NVIDIA training stack. Windows users should choose a supported native stack or a documented WSL2 configuration rather than mixing Linux commands into an unrelated Windows environment. Exact support depends on the packages and GPU; consult the official installation instructions for the selected release. Do not install a random wheel merely because a forum comment says it fixes an error.
A command may keep running for minutes or hours. Output scrolling past is a log, not proof of progress. Check the process, step counter, throughput, device utilization, and written checkpoints. Ctrl+C requests interruption; whether a useful checkpoint remains depends on the program. Learn recovery on a tiny run before relying on it for an overnight job.
Python files and environments
Python reads instructions from a .py file. Indentation is meaningful. Copying code from a formatted page can introduce broken quotation marks or indentation, which is one reason the companion files matter. A comment begins with # and is not executed. A string is quoted text. An integer is a whole number. A float stores an approximate real number. A list holds an ordered sequence. A dictionary maps keys to values.
learning_rate = 0.001
class_names = ["refund", "shipping", "other"]
record = {"text": "Where is my package?", "label": "shipping"}
print(record["text"])
Here learning_rate is a variable. The square brackets after record retrieve the value stored under the key text. A function packages a repeatable computation; calling it runs that computation. A class can combine data and behavior, as a neural-network class combines trainable parameters and a forward method.
Libraries such as PyTorch provide large collections of existing functions. A Python environment determines which versions of those libraries are available. Installing a package into one environment and launching Python from another is a common source of errors. Prefer python -m pip to a bare pip command, because it identifies which interpreter should manage the packages.
A virtual environment is an isolated package directory. A typical project starts like this:
python -m venv .venv
source .venv/bin/activate
python --version
python -m pip --version
On PowerShell the activation command differs. Activation is convenience, not magic: you can invoke the environment's Python by its full path. Record the Python version. Then follow the project-specific dependency file and official PyTorch wheel selection. CUDA support in PyTorch must be compatible with the installed driver and GPU. A CUDA toolkit installed somewhere on the machine is not proof that the Python process has a CUDA-enabled PyTorch build.
Do not combine all projects in this book into one enormous environment. The language, diffusion, audio, and classical-ML projects may need different versions. Create one environment per project, save its package list after the smoke test, and change one dependency at a time when diagnosing a failure.
Check the environment before the model
On an NVIDIA machine, nvidia-smi reports the GPU, driver, running processes, and memory use. It is a diagnostic tool, not a training benchmark. Once PyTorch is installed, run this small check in the same environment that will train:
import torch
print("PyTorch:", torch.__version__)
print("CUDA available:", torch.cuda.is_available())
if torch.cuda.is_available():
print("GPU:", torch.cuda.get_device_name(0))
print("BF16 supported:", torch.cuda.is_bf16_supported())
print("VRAM bytes:", torch.cuda.get_device_properties(0).total_memory)
CUDA is NVIDIA's GPU computing platform. FP32, FP16, and BF16 are numeric formats described in the memory chapter. Do not infer BF16 support from the phrase RTX alone. Check the actual device and framework support. A recipe that silently falls back to CPU can appear to work while taking far longer than expected; the training log should print the selected device.
A GPU test that only allocates a tensor verifies less than a training smoke test. A useful smoke test performs a forward pass, computes a finite loss, runs backward, updates parameters, saves a checkpoint, loads it, and performs inference. It should finish quickly and leave behind evidence you can inspect.
A workspace that you can reason about
Use a structure that makes raw material and derived artifacts easy to distinguish:
project/
README.md
requirements.txt
configs/
data/
raw/
processed/
splits/
src/
runs/
experiment_001/
evaluation/
exports/
Keep raw data unchanged. Write cleaned or tokenized versions elsewhere. Give each run its own directory. Never point a new experiment at an old output directory unless the code explicitly supports resuming and you intend to resume. A useful run directory contains the resolved configuration, data hashes, model revision, log, evaluation outputs, and checkpoint metadata.
Use version control for code, small configurations, and dataset-generation rules. Large model files and private datasets usually belong in a separate artifact store rather than ordinary Git history. Do not commit secrets. A token used to download a gated model is a credential, not part of the experiment. If logs accidentally contain a secret, treat it as exposed and follow the service's revocation process.
Exercises
Create a new directory and run a Python file that prints a sentence. Make a deliberate typo, read the error, then correct it. Locate the file from the terminal rather than relying only on an editor's open tab.
Write a JSON record containing an input, a desired output, and a source identifier. Reopen it with Python's json module. Print the desired output. Then explain the difference between the file's contents, the Python object loaded from it, and a tensor that a model would consume.
Before any GPU training, write down the exact GPU model, available VRAM, system RAM, free disk space, driver version, Python version, and PyTorch version. If one entry is unknown, the preparation is not finished.