Training Your
Own Models
Browse the book
Reference16 min read

Glossary with links to explanations

243 terms

Use each link to return to the explanation and its worked project. Terms introduced at several levels point to a useful first or central explanation.

  • Activation. An intermediate value produced during a forward pass. Explanation
  • Activation checkpointing. Recomputing selected activations during backward to save memory. Explanation
  • Activation function. A nonlinear transformation such as ReLU or GELU, distinct from an intermediate activation value. Explanation
  • Active parameters. The subset involved in processing a token; not the complete resident or trainable parameter count. Explanation
  • AdamW. An adaptive optimizer with decoupled weight decay. Explanation
  • Approximate nearest neighbors. An indexing approach trading some exact-search fidelity for faster vector retrieval. Explanation
  • Architecture. The form of the model computation and arrangement of its learned parts. Explanation
  • Architecture/data relationship. The connection between what a model can represent and what its examples must demonstrate. Explanation
  • ASR hallucination. Text generated by a recognizer despite the corresponding words being absent from the audio; silence and noise probes help expose this. Explanation
  • Assistant-only loss. Loss on assistant responses, while user, system and tool-observation text remains context. Explanation
  • Attention. A learned way for a position to combine information from other permitted positions. Explanation
  • Attention mask. A rule controlling which positions may influence a representation. Explanation
  • Audio-transcript alignment. Ensuring that a segment’s target text contains exactly the words present in that segment, rather than the transcript of a longer recording. Explanation
  • Autocast. A framework policy choosing compute precision for supported operations without necessarily changing stored parameter dtype. Explanation
  • Autograd. Automatic differentiation used to calculate gradients. Explanation
  • Automatic speech recognition (ASR). Converting recorded speech audio into a written transcript; distinct from generating speech, identifying speakers, or summarizing a transcript. Explanation
  • Average precision. A ranking metric summarizing precision as recall increases. Explanation
  • Backpropagation. Computing parameter gradients through the model computation. Explanation
  • Base model. A pretrained model before a particular instruction or task adaptation. Explanation
  • Baseline. A simpler or unchanged system used as the comparison before training. Explanation
  • Batch. Examples processed together or contributing to one update, depending on the stated unit. Explanation
  • BCE. Binary cross-entropy, a loss for binary targets such as foreground versus background. Explanation
  • BF16. A two-byte floating-point format with a different precision/range tradeoff from FP16. Explanation
  • BF16 master-state distinction. The difference between low-precision operation dtype and higher-precision stored parameters or optimizer state. Explanation
  • Bias. An additive learned value in a model transformation. Explanation
  • Boundary F1. Precision/recall-based agreement of boundaries under a specified spatial tolerance. Explanation
  • Boundary tolerance. The distance within which predicted and reference boundary points count as matching. Explanation
  • Calibration. Checking and adjusting the relationship between predicted probabilities and observed outcomes. Explanation
  • Calibration split. Held-out examples used to fit a probability correction without fitting the original predictor. Explanation
  • Catastrophic forgetting. Loss of previously useful behavior during adaptation to a different or narrower distribution. Explanation
  • Causal mask. A mask that prevents a position from reading future positions. Explanation
  • Character error rate (CER). The analogous edit rate over declared character units; useful when whitespace-delimited words are unsuitable. Explanation
  • Chat template. A model-specific serialization that converts roles, turns and optional tools into the tokens expected by a chat checkpoint. Explanation
  • Checkpoint. Saved model state, with additional optimizer and random state when intended for resume. Explanation
  • Class imbalance. Unequal representation of target classes, often requiring slice-aware evaluation. Explanation
  • Classifier free guidance. A sampling technique combining conditional and unconditional predictions. Explanation
  • CLM. A contrastive approach comparing separately encoded state and action representations. Explanation
  • Closed-loop evaluation. Testing trajectories using the model’s own calls and the observations they actually produce. Explanation
  • Code-switching. Alternating languages within speech, which requires separate evaluation rather than assuming single-language quality transfers. Explanation
  • Completion-only loss. Loss computed on the completion after a prompt, excluding prompt targets. Explanation
  • Conformer. An acoustic architecture combining local convolutional processing with attention over a longer context. Explanation
  • Connectionist temporal classification (CTC). A sequence objective summing probabilities of frame-level paths that collapse to the desired transcript, including a blank label. Explanation
  • Context length. The amount of input history available to a sequence model. Explanation
  • Continued pretraining. Continuing a pretrained model’s original language-modeling objective on additional text. Explanation
  • Contour. A boundary extracted from a predicted or labeled region. Explanation
  • Contour hierarchy. Relationships among outer boundaries and nested holes or regions. Explanation
  • Contrastive loss. An objective making matched representations score above selected alternatives. Explanation
  • Convolution. A learned local filter reused across spatial or temporal positions. Explanation
  • Cosine similarity. The normalized dot product measuring vector direction similarity. Explanation
  • CPU dependency setup. Installing and verifying the packages required by a CPU execution path in an isolated environment. Explanation
  • CPU/GPU inference. Running the same learned function on different devices with compatible preprocessing, numeric policy, and output checks. Explanation
  • Cross entropy. A loss penalizing low probability assigned to the correct target. Explanation
  • CTC blank. A special no-output label used when collapsing CTC paths; distinct from a written space. Explanation
  • Data augmentation. Training-input transformations justified by a target-preservation assumption. Explanation
  • Data leakage. Information crossing an evaluation boundary or appearing before it would be available. Explanation
  • Dataset card. A record of data sources, rights, processing, splits, coverage, and limitations. Explanation
  • Decode. Producing subsequent tokens incrementally. Explanation
  • Dense model. A model whose dense blocks normally participate for each token, unlike selectively routed experts. Explanation
  • Depthwise convolution. Local filtering independently within channels. Explanation
  • Detection. Predicting object locations and usually their classes. Explanation
  • Device budget. The combined memory, compute, storage, power, and response-time constraints of the target device. Explanation
  • Dice loss. An overlap-based objective derived from the Dice similarity of predicted and target regions. Explanation
  • Diffusion. A generative approach learning to reverse a noise-corruption process. Explanation
  • Distillation. Training a student using selected teacher behavior or distributions. Explanation
  • Double quantization. Quantizing quantization constants as well as base weights to reduce storage overhead. Explanation
  • DPO. Direct Preference Optimization: fitting chosen-versus-rejected preferences relative to a reference policy. Explanation
  • Dropout. Random suppression during training that must be disabled for ordinary evaluation. Explanation
  • Dual encoder. An architecture encoding the two sides of a comparison separately. Explanation
  • Effective batch. All microbatches contributing to one optimizer update. Explanation
  • Embedding. A learned vector representation used for a specified prediction, similarity, or retrieval task. Explanation
  • Embedding table. A learned lookup from discrete IDs to vectors. Explanation
  • Encoder-decoder. A model that forms an intermediate representation and transforms it into an output representation. Explanation
  • Epoch. One pass through a defined dataset, where that concept applies. Explanation
  • Evaluation. A fixed procedure for comparing task behavior against requirements. Explanation
  • Expert parallelism. Placing experts on separate devices and dispatching tokens. Explanation
  • Export. Packaging a learned artifact and its interface for a particular inference runtime. Explanation
  • F1. The harmonic mean of precision and recall. Explanation
  • False negative. A genuinely positive or relevant example treated as negative by a decision or training construction. Explanation
  • Feature. Information represented in a form the model can consume. Explanation
  • Feature channel. One learned component of a representation at each spatial or temporal position. Explanation
  • Final holdout protocol. Evaluating a protected test set only after model and pipeline choices are fixed. Explanation
  • Fine tuning. Further training starting from pretrained weights. Explanation
  • FlashAttention. IO-aware execution of exact attention. Explanation
  • Flow matching. Learning a vector field for a path between noise and data distributions. Explanation
  • Forced alignment. Estimating where supplied transcript units occur in an audio recording; it does not verify that the supplied words are correct. Explanation
  • Foreground imbalance. A segmentation condition where target-object pixels are much rarer than background pixels. Explanation
  • FP16. A two-byte floating-point format whose numeric behavior differs from BF16. Explanation
  • FP32. A four-byte floating-point format. Explanation
  • Fresh environment. An isolated package environment that avoids accidental dependency inheritance from another project. Explanation
  • Frozen feature extractor. A pretrained component whose parameters remain unchanged while another part learns. Explanation
  • Full fine-tuning. Updating all model parameters starting from pretrained weights; it does not mean random initialization. Explanation
  • Full parameter training. Updating every selected model parameter rather than only an adapter or head. Explanation
  • GQA. Grouped-query attention: several query heads share a smaller number of key/value heads. Explanation
  • GQA and MQA. Sharing KV heads among several or all query heads. Explanation
  • Gradient. The local rate of change of loss with respect to a parameter. Explanation
  • Gradient accumulation. Combining several microbatch gradients before an optimizer update. Explanation
  • Gradient boosting. An ensemble built by adding learners that improve the remaining prediction error. Explanation
  • Gradient caching. A technique recomputing or caching parts of embedding training to support a larger logical contrastive batch. Explanation
  • Gradient checkpointing. Recomputing selected forward activations during backward to trade extra compute for lower memory use. Explanation
  • Gradient clipping. Limiting unusually large gradients before updating. Explanation
  • Grouped split. Keeping related records in the same partition to protect evaluation independence. Explanation
  • Hard negative. An incorrect candidate that is plausibly confusable with the correct match. Explanation
  • Hybrid. Combination of different sequence-mixing mechanisms. Explanation
  • Hyperparameter. A configuration choice controlling the model or training procedure. Explanation
  • Imputation. Filling missing values using a rule fitted only on appropriate training data. Explanation
  • In-batch negatives. Other examples in a training batch used as alternative candidates in a contrastive objective. Explanation
  • Index consistency. Agreement among the saved encoder, preprocessing, vector representation, and indexed documents. Explanation
  • Inference. Using fitted model parameters to produce an output. Explanation
  • Input. The information available to the model when a prediction is made. Explanation
  • Instance segmentation. Predicting distinct masks for individual objects. Explanation
  • IoU. Intersection over union, measuring region overlap relative to the combined region. Explanation
  • Jev. A hosted typed-decision model; its private internals are not reproduced here. Explanation
  • KDA. Kimi Delta Attention, a gated recurrent linear-attention mechanism. Explanation
  • Kev. An open decision-model implementation studied separately from Jev. Explanation
  • Knowledge editing. Methods designed to change targeted factual associations, with locality/generalization limits requiring evaluation. Explanation
  • KV cache. Stored attention state used to avoid repeated generation computation. Explanation
  • L2 normalization. Rescaling a vector to unit Euclidean length. Explanation
  • Label. The recorded correct target for an example. Explanation
  • Last-token pooling. Using the final nonpadding token representation as the sequence embedding under the model's trained contract. Explanation
  • Latent autoencoder. A model mapping images or audio to and from a smaller learned representation. Explanation
  • Laya. A public ModernBERT-based typed-decision model used as a study case. Explanation
  • Leakage. Information crossing an evaluation boundary or becoming available earlier than it would at prediction time. Explanation
  • Learnability. Whether the desired mapping can be inferred from the available information and examples under the chosen model. Explanation
  • Learning curve. A comparison of performance as training data or training budget increases. Explanation
  • Learning rate. The scale of an optimizer update. Explanation
  • Linear attention. Sequence mixing that avoids the usual quadratic pair matrix. Explanation
  • Load balancing. Mechanism to avoid unusably uneven expert traffic. Explanation
  • Log-Mel spectrogram. Successive short-window frequency representations pooled through Mel filters and logarithmically scaled, using the model processor’s expected conventions. Explanation
  • Logistic regression. A linear score transformed into a class probability, fitted from labeled examples. Explanation
  • Logit. An unnormalized prediction score before probability conversion. Explanation
  • Logits. Unnormalized scores over the vocabulary before conversion to token probabilities. Explanation
  • LoRA. Low-rank adaptation: freeze a base matrix and learn a smaller factorized correction. Explanation
  • LoRA alpha. A scale parameter; standard LoRA multiplies the adapter correction by alpha divided by rank. Explanation
  • LoRA rank. The inner dimension of the two adapter matrices, controlling the rank and capacity of their product. Explanation
  • Loss. A numeric training objective comparing prediction and target. Explanation
  • Loss mask. A rule selecting which predictions contribute to training loss. Explanation
  • Macro-F1. The average of per-class F1 scores, giving each class equal weight. Explanation
  • MAE. Mean absolute error, measured in the target's units. Explanation
  • Margin. A required score or distance separation between desired and undesired outcomes. Explanation
  • Master weights. An additional high-precision parameter copy maintained by some mixed-precision training schemes. Explanation
  • Mean pooling. Averaging selected position representations into one vector. Explanation
  • mHC. Constrained learned routing among residual streams. Explanation
  • Microbatch. The examples handled in one forward/backward pass. Explanation
  • Mixed precision. Using different numeric formats for different parts of training. Explanation
  • Mixture of experts. A model that routes each token through selected expert networks while retaining the larger expert collection. Explanation
  • MLA. Learned latent representation of attention KV information. Explanation
  • Model card. A record of model purpose, provenance, evaluation, tested targets, and limitations. Explanation
  • Model/index SHA contract. Using cryptographic file identities to ensure a saved encoder, preprocessing, and vector index belong together. Explanation
  • MRR. Mean reciprocal rank of the first relevant result. Explanation
  • MTP. Training prediction targets beyond the immediate next token. Explanation
  • Multimodal embeddings. Representations that support matching or retrieval across different input modalities. Explanation
  • Multimodal input. Input combining more than one type of information, such as text and audio. Explanation
  • nDCG. Normalized discounted cumulative gain, a ranking metric rewarding relevant results near the top. Explanation
  • NF4. NormalFloat four-bit quantization, designed for representing approximately normally distributed weights. Explanation
  • On-policy distillation. Teacher supervision on the student's own continuations. Explanation
  • Optimizer. The rule that updates parameters using gradients and possibly history. Explanation
  • Optimizer state. Persistent tensors an optimizer maintains beyond weights and gradients, such as Adam first and second moments. Explanation
  • Option pointer. A readout scoring candidate options supplied in the input. Explanation
  • Out-of-distribution. Inputs that differ meaningfully from the conditions represented in training and evaluation. Explanation
  • Overfitting. Improving fit to training examples without matching improvement on new cases. Explanation
  • Padding. Placeholder positions added to make sequence shapes compatible. Explanation
  • Parameter. An adjustable learned value in the model. Explanation
  • Parameter sharding. Dividing stored training parameters among devices. Explanation
  • Partial fine-tuning. Updating selected existing parameters while freezing the others; the ASR starter freezes its acoustic tower and does not use LoRA. Explanation
  • Patch merger. Combination of neighboring features into fewer input units. Explanation
  • Pipeline parallelism. Distributing successive layer groups across stages. Explanation
  • Polygon simplification. Reducing boundary vertices while bounding an allowed geometric change. Explanation
  • Pooling. Combining several positions into a smaller summary representation. Explanation
  • Precision. The fraction of predicted positive cases that are actually positive. Explanation
  • Prediction contract. A specification of input fields, target meaning, output format, timing, and failure handling. Explanation
  • Prefill. Processing the supplied prefix or prompt. Explanation
  • Pretraining. Learning an initial representation or generative model from a broad objective. Explanation
  • Pretraining from scratch. Training a newly initialized model rather than adapting pretrained weights. Explanation
  • Probability calibration. Agreement between stated probabilities and observed outcome frequencies. Explanation
  • Projector. Map from encoder features to a backbone's feature width. Explanation
  • QAT. Training that accounts for quantization effects. Explanation
  • QLoRA. Training floating-point low-rank adapters through a frozen quantized language-model base. Explanation
  • Quantization. Representing selected numeric values with fewer bits under a specific runtime format. Explanation
  • Query instruction. A task description added to a retrieval query in the format expected by an instruction-aware embedding model. Explanation
  • Qwen3 embedding representation. The selected Qwen encoder's last-token-pooled, normalized vector with model-specific query formatting. Explanation
  • R-squared. A regression comparison against predicting the mean, whose interpretation depends on the evaluation data. Explanation
  • RAG. Retrieval-augmented generation: supplying retrieved external evidence to a model at inference time. Explanation
  • Recall. The fraction of actual positive cases that the system finds. Explanation
  • Recall@k. A retrieval measure of relevant results found among the first k candidates under the task's relevance definition. Explanation
  • Recurrent state. Persistent summary updated as tokens arrive. Explanation
  • Reference policy. A fixed baseline distribution used to measure how the trained policy changes completion likelihoods. Explanation
  • Regression. Predicting a numeric quantity rather than a discrete class. Explanation
  • Regularization. Constraints or training choices intended to improve generalization rather than only training fit. Explanation
  • Reliability. In calibration, the relationship between predicted probability bins and observed outcomes. Explanation
  • Reranker. A second-stage model scoring a smaller candidate set more precisely. Explanation
  • Resampling. Changing an audio sample grid with suitable filtering; downsampling must suppress frequencies above the new Nyquist limit. Explanation
  • Resident parameters. Complete weights that must be stored somewhere. Explanation
  • Residual connection. Adding an earlier representation to a learned transformation. Explanation
  • RMSE. Root mean squared error, emphasizing larger errors while retaining target units. Explanation
  • ROC-AUC. A score-ranking measure comparing positive and negative examples across thresholds. Explanation
  • RoPE. Rotary positional embeddings: positional relationships represented by rotations in attention coordinates. Explanation
  • Router. Learned function that selects experts for token representations. Explanation
  • Run manifest. A saved record of model/data identity, configuration, versions and execution details for one experiment. Explanation
  • RVQ. Several codebooks used successively for discrete signal representation. Explanation
  • Sample rate. The number of waveform amplitude measurements per second; changing a file header is not resampling. Explanation
  • Seed. An initial state for pseudo-random generation. Explanation
  • Semantic segmentation. Assigning a class to each pixel without necessarily separating individual instances. Explanation
  • Sequence distillation. Using teacher-produced complete output sequences as student training targets. Explanation
  • SFT. Supervised fine-tuning on examples of desired output conditioned on the supplied input. Explanation
  • Shared expert. Expert used for all tokens alongside selected experts. Explanation
  • Short convolution. A sequence-mixing operation over a local window, used with input-dependent gates in LFM2. Explanation
  • Skip connection. A route carrying earlier features to a later stage, often preserving detail. Explanation
  • Sparse attention. Attention over a selected subset of positions. Explanation
  • Speaker diarization. Estimating which speaker spoke when; a separate task from recognizing the spoken words. Explanation
  • Speculative decoding. Draft proposals checked by a target model. Explanation
  • SSM. Structured state-space sequence model. Explanation
  • Structured output. An output governed by explicit fields, types, labels, or other schema constraints. Explanation
  • SWA. Direct attention limited to a moving local window. Explanation
  • Target. The desired output used to define correctness during training. Explanation
  • Target modules. The specific network modules selected to receive trainable adapters. Explanation
  • Target scaling. Transforming target values during training and reversing the transformation for output. Explanation
  • Teacher forcing. Training predictions on a supplied correct history, rather than the model’s own earlier generated mistakes. Explanation
  • Temperature in calibration. A held-out-fitted logit scale used to improve probability calibration. Explanation
  • Temperature in contrastive learning. A scale controlling the sharpness of contrastive candidate probabilities. Explanation
  • Temperature in generation. A logit rescaling used to change a sampled token distribution. Explanation
  • Tensor. A numeric array with shape, dtype, and device. Explanation
  • Tensor parallelism. Splitting matrix operations across devices. Explanation
  • Test. A held-out evaluation role used after development choices are fixed. Explanation
  • Test set. Held-out data used after development choices are fixed. Explanation
  • TF-IDF. Term frequency-inverse document frequency, a sparse representation weighting local term counts against corpus prevalence. Explanation
  • Threshold. A selected cutoff converting a score into a decision. Explanation
  • Tied embeddings. Reusing one parameter matrix for token input embeddings and vocabulary output projection. Explanation
  • TinyML. Machine learning designed around tight device memory, compute, power, and latency limits. Explanation
  • Tokenization. Converting text into the discrete token identifiers expected by a particular model. Explanation
  • Tokenizer. A mapping between text and model token identifiers. Explanation
  • Tool observation. The result returned by the external executor and supplied as context to the model. Explanation
  • Tool schema. A structured specification of a tool name, purpose, argument types and constraints. Explanation
  • Top-k routing. Selecting k expert assignments for each token. Explanation
  • Training from scratch. Optimizing randomly initialized parameters rather than adapting pretrained weights. Explanation
  • Transfer learning. Reusing a pretrained representation for a new task. Explanation
  • Triplet loss. An objective comparing an anchor, a positive match, and a negative example. Explanation
  • Truncation. Discarding input or target content beyond a length limit. Explanation
  • Validation. An evaluation role used to compare candidates and select configuration or checkpoints. Explanation
  • Validation set. Held-out data used to guide experiment choices. Explanation
  • Virtual environment. An isolated Python package directory tied to a chosen interpreter. Explanation
  • Voice activity detection (VAD). Predicting which audio regions likely contain speech, with a tradeoff between non-speech suppression and accidentally deleting quiet speech. Explanation
  • Weight decay. A parameter-shrinking term applied by the optimizer. Explanation
  • Word error rate (WER). Substitutions plus deletions plus insertions, divided by the number of reference words under a specified normalization/tokenization policy. Explanation