Glossary with links to explanations
243 terms
Use each link to return to the explanation and its worked project. Terms introduced at several levels point to a useful first or central explanation.
- Activation. An intermediate value produced during a forward pass. Explanation
- Activation checkpointing. Recomputing selected activations during backward to save memory. Explanation
- Activation function. A nonlinear transformation such as ReLU or GELU, distinct from an intermediate activation value. Explanation
- Active parameters. The subset involved in processing a token; not the complete resident or trainable parameter count. Explanation
- AdamW. An adaptive optimizer with decoupled weight decay. Explanation
- Approximate nearest neighbors. An indexing approach trading some exact-search fidelity for faster vector retrieval. Explanation
- Architecture. The form of the model computation and arrangement of its learned parts. Explanation
- Architecture/data relationship. The connection between what a model can represent and what its examples must demonstrate. Explanation
- ASR hallucination. Text generated by a recognizer despite the corresponding words being absent from the audio; silence and noise probes help expose this. Explanation
- Assistant-only loss. Loss on assistant responses, while user, system and tool-observation text remains context. Explanation
- Attention. A learned way for a position to combine information from other permitted positions. Explanation
- Attention mask. A rule controlling which positions may influence a representation. Explanation
- Audio-transcript alignment. Ensuring that a segment’s target text contains exactly the words present in that segment, rather than the transcript of a longer recording. Explanation
- Autocast. A framework policy choosing compute precision for supported operations without necessarily changing stored parameter dtype. Explanation
- Autograd. Automatic differentiation used to calculate gradients. Explanation
- Automatic speech recognition (ASR). Converting recorded speech audio into a written transcript; distinct from generating speech, identifying speakers, or summarizing a transcript. Explanation
- Average precision. A ranking metric summarizing precision as recall increases. Explanation
- Backpropagation. Computing parameter gradients through the model computation. Explanation
- Base model. A pretrained model before a particular instruction or task adaptation. Explanation
- Baseline. A simpler or unchanged system used as the comparison before training. Explanation
- Batch. Examples processed together or contributing to one update, depending on the stated unit. Explanation
- BCE. Binary cross-entropy, a loss for binary targets such as foreground versus background. Explanation
- BF16. A two-byte floating-point format with a different precision/range tradeoff from FP16. Explanation
- BF16 master-state distinction. The difference between low-precision operation dtype and higher-precision stored parameters or optimizer state. Explanation
- Bias. An additive learned value in a model transformation. Explanation
- Boundary F1. Precision/recall-based agreement of boundaries under a specified spatial tolerance. Explanation
- Boundary tolerance. The distance within which predicted and reference boundary points count as matching. Explanation
- Calibration. Checking and adjusting the relationship between predicted probabilities and observed outcomes. Explanation
- Calibration split. Held-out examples used to fit a probability correction without fitting the original predictor. Explanation
- Catastrophic forgetting. Loss of previously useful behavior during adaptation to a different or narrower distribution. Explanation
- Causal mask. A mask that prevents a position from reading future positions. Explanation
- Character error rate (CER). The analogous edit rate over declared character units; useful when whitespace-delimited words are unsuitable. Explanation
- Chat template. A model-specific serialization that converts roles, turns and optional tools into the tokens expected by a chat checkpoint. Explanation
- Checkpoint. Saved model state, with additional optimizer and random state when intended for resume. Explanation
- Class imbalance. Unequal representation of target classes, often requiring slice-aware evaluation. Explanation
- Classifier free guidance. A sampling technique combining conditional and unconditional predictions. Explanation
- CLM. A contrastive approach comparing separately encoded state and action representations. Explanation
- Closed-loop evaluation. Testing trajectories using the model’s own calls and the observations they actually produce. Explanation
- Code-switching. Alternating languages within speech, which requires separate evaluation rather than assuming single-language quality transfers. Explanation
- Completion-only loss. Loss computed on the completion after a prompt, excluding prompt targets. Explanation
- Conformer. An acoustic architecture combining local convolutional processing with attention over a longer context. Explanation
- Connectionist temporal classification (CTC). A sequence objective summing probabilities of frame-level paths that collapse to the desired transcript, including a blank label. Explanation
- Context length. The amount of input history available to a sequence model. Explanation
- Continued pretraining. Continuing a pretrained model’s original language-modeling objective on additional text. Explanation
- Contour. A boundary extracted from a predicted or labeled region. Explanation
- Contour hierarchy. Relationships among outer boundaries and nested holes or regions. Explanation
- Contrastive loss. An objective making matched representations score above selected alternatives. Explanation
- Convolution. A learned local filter reused across spatial or temporal positions. Explanation
- Cosine similarity. The normalized dot product measuring vector direction similarity. Explanation
- CPU dependency setup. Installing and verifying the packages required by a CPU execution path in an isolated environment. Explanation
- CPU/GPU inference. Running the same learned function on different devices with compatible preprocessing, numeric policy, and output checks. Explanation
- Cross entropy. A loss penalizing low probability assigned to the correct target. Explanation
- CTC blank. A special no-output label used when collapsing CTC paths; distinct from a written space. Explanation
- Data augmentation. Training-input transformations justified by a target-preservation assumption. Explanation
- Data leakage. Information crossing an evaluation boundary or appearing before it would be available. Explanation
- Dataset card. A record of data sources, rights, processing, splits, coverage, and limitations. Explanation
- Decode. Producing subsequent tokens incrementally. Explanation
- Dense model. A model whose dense blocks normally participate for each token, unlike selectively routed experts. Explanation
- Depthwise convolution. Local filtering independently within channels. Explanation
- Detection. Predicting object locations and usually their classes. Explanation
- Device budget. The combined memory, compute, storage, power, and response-time constraints of the target device. Explanation
- Dice loss. An overlap-based objective derived from the Dice similarity of predicted and target regions. Explanation
- Diffusion. A generative approach learning to reverse a noise-corruption process. Explanation
- Distillation. Training a student using selected teacher behavior or distributions. Explanation
- Double quantization. Quantizing quantization constants as well as base weights to reduce storage overhead. Explanation
- DPO. Direct Preference Optimization: fitting chosen-versus-rejected preferences relative to a reference policy. Explanation
- Dropout. Random suppression during training that must be disabled for ordinary evaluation. Explanation
- Dual encoder. An architecture encoding the two sides of a comparison separately. Explanation
- Effective batch. All microbatches contributing to one optimizer update. Explanation
- Embedding. A learned vector representation used for a specified prediction, similarity, or retrieval task. Explanation
- Embedding table. A learned lookup from discrete IDs to vectors. Explanation
- Encoder-decoder. A model that forms an intermediate representation and transforms it into an output representation. Explanation
- Epoch. One pass through a defined dataset, where that concept applies. Explanation
- Evaluation. A fixed procedure for comparing task behavior against requirements. Explanation
- Expert parallelism. Placing experts on separate devices and dispatching tokens. Explanation
- Export. Packaging a learned artifact and its interface for a particular inference runtime. Explanation
- F1. The harmonic mean of precision and recall. Explanation
- False negative. A genuinely positive or relevant example treated as negative by a decision or training construction. Explanation
- Feature. Information represented in a form the model can consume. Explanation
- Feature channel. One learned component of a representation at each spatial or temporal position. Explanation
- Final holdout protocol. Evaluating a protected test set only after model and pipeline choices are fixed. Explanation
- Fine tuning. Further training starting from pretrained weights. Explanation
- FlashAttention. IO-aware execution of exact attention. Explanation
- Flow matching. Learning a vector field for a path between noise and data distributions. Explanation
- Forced alignment. Estimating where supplied transcript units occur in an audio recording; it does not verify that the supplied words are correct. Explanation
- Foreground imbalance. A segmentation condition where target-object pixels are much rarer than background pixels. Explanation
- FP16. A two-byte floating-point format whose numeric behavior differs from BF16. Explanation
- FP32. A four-byte floating-point format. Explanation
- Fresh environment. An isolated package environment that avoids accidental dependency inheritance from another project. Explanation
- Frozen feature extractor. A pretrained component whose parameters remain unchanged while another part learns. Explanation
- Full fine-tuning. Updating all model parameters starting from pretrained weights; it does not mean random initialization. Explanation
- Full parameter training. Updating every selected model parameter rather than only an adapter or head. Explanation
- GQA. Grouped-query attention: several query heads share a smaller number of key/value heads. Explanation
- GQA and MQA. Sharing KV heads among several or all query heads. Explanation
- Gradient. The local rate of change of loss with respect to a parameter. Explanation
- Gradient accumulation. Combining several microbatch gradients before an optimizer update. Explanation
- Gradient boosting. An ensemble built by adding learners that improve the remaining prediction error. Explanation
- Gradient caching. A technique recomputing or caching parts of embedding training to support a larger logical contrastive batch. Explanation
- Gradient checkpointing. Recomputing selected forward activations during backward to trade extra compute for lower memory use. Explanation
- Gradient clipping. Limiting unusually large gradients before updating. Explanation
- Grouped split. Keeping related records in the same partition to protect evaluation independence. Explanation
- Hard negative. An incorrect candidate that is plausibly confusable with the correct match. Explanation
- Hybrid. Combination of different sequence-mixing mechanisms. Explanation
- Hyperparameter. A configuration choice controlling the model or training procedure. Explanation
- Imputation. Filling missing values using a rule fitted only on appropriate training data. Explanation
- In-batch negatives. Other examples in a training batch used as alternative candidates in a contrastive objective. Explanation
- Index consistency. Agreement among the saved encoder, preprocessing, vector representation, and indexed documents. Explanation
- Inference. Using fitted model parameters to produce an output. Explanation
- Input. The information available to the model when a prediction is made. Explanation
- Instance segmentation. Predicting distinct masks for individual objects. Explanation
- IoU. Intersection over union, measuring region overlap relative to the combined region. Explanation
- Jev. A hosted typed-decision model; its private internals are not reproduced here. Explanation
- KDA. Kimi Delta Attention, a gated recurrent linear-attention mechanism. Explanation
- Kev. An open decision-model implementation studied separately from Jev. Explanation
- Knowledge editing. Methods designed to change targeted factual associations, with locality/generalization limits requiring evaluation. Explanation
- KV cache. Stored attention state used to avoid repeated generation computation. Explanation
- L2 normalization. Rescaling a vector to unit Euclidean length. Explanation
- Label. The recorded correct target for an example. Explanation
- Last-token pooling. Using the final nonpadding token representation as the sequence embedding under the model's trained contract. Explanation
- Latent autoencoder. A model mapping images or audio to and from a smaller learned representation. Explanation
- Laya. A public ModernBERT-based typed-decision model used as a study case. Explanation
- Leakage. Information crossing an evaluation boundary or becoming available earlier than it would at prediction time. Explanation
- Learnability. Whether the desired mapping can be inferred from the available information and examples under the chosen model. Explanation
- Learning curve. A comparison of performance as training data or training budget increases. Explanation
- Learning rate. The scale of an optimizer update. Explanation
- Linear attention. Sequence mixing that avoids the usual quadratic pair matrix. Explanation
- Load balancing. Mechanism to avoid unusably uneven expert traffic. Explanation
- Log-Mel spectrogram. Successive short-window frequency representations pooled through Mel filters and logarithmically scaled, using the model processor’s expected conventions. Explanation
- Logistic regression. A linear score transformed into a class probability, fitted from labeled examples. Explanation
- Logit. An unnormalized prediction score before probability conversion. Explanation
- Logits. Unnormalized scores over the vocabulary before conversion to token probabilities. Explanation
- LoRA. Low-rank adaptation: freeze a base matrix and learn a smaller factorized correction. Explanation
- LoRA alpha. A scale parameter; standard LoRA multiplies the adapter correction by alpha divided by rank. Explanation
- LoRA rank. The inner dimension of the two adapter matrices, controlling the rank and capacity of their product. Explanation
- Loss. A numeric training objective comparing prediction and target. Explanation
- Loss mask. A rule selecting which predictions contribute to training loss. Explanation
- Macro-F1. The average of per-class F1 scores, giving each class equal weight. Explanation
- MAE. Mean absolute error, measured in the target's units. Explanation
- Margin. A required score or distance separation between desired and undesired outcomes. Explanation
- Master weights. An additional high-precision parameter copy maintained by some mixed-precision training schemes. Explanation
- Mean pooling. Averaging selected position representations into one vector. Explanation
- mHC. Constrained learned routing among residual streams. Explanation
- Microbatch. The examples handled in one forward/backward pass. Explanation
- Mixed precision. Using different numeric formats for different parts of training. Explanation
- Mixture of experts. A model that routes each token through selected expert networks while retaining the larger expert collection. Explanation
- MLA. Learned latent representation of attention KV information. Explanation
- Model card. A record of model purpose, provenance, evaluation, tested targets, and limitations. Explanation
- Model/index SHA contract. Using cryptographic file identities to ensure a saved encoder, preprocessing, and vector index belong together. Explanation
- MRR. Mean reciprocal rank of the first relevant result. Explanation
- MTP. Training prediction targets beyond the immediate next token. Explanation
- Multimodal embeddings. Representations that support matching or retrieval across different input modalities. Explanation
- Multimodal input. Input combining more than one type of information, such as text and audio. Explanation
- nDCG. Normalized discounted cumulative gain, a ranking metric rewarding relevant results near the top. Explanation
- NF4. NormalFloat four-bit quantization, designed for representing approximately normally distributed weights. Explanation
- On-policy distillation. Teacher supervision on the student's own continuations. Explanation
- Optimizer. The rule that updates parameters using gradients and possibly history. Explanation
- Optimizer state. Persistent tensors an optimizer maintains beyond weights and gradients, such as Adam first and second moments. Explanation
- Option pointer. A readout scoring candidate options supplied in the input. Explanation
- Out-of-distribution. Inputs that differ meaningfully from the conditions represented in training and evaluation. Explanation
- Overfitting. Improving fit to training examples without matching improvement on new cases. Explanation
- Padding. Placeholder positions added to make sequence shapes compatible. Explanation
- Parameter. An adjustable learned value in the model. Explanation
- Parameter sharding. Dividing stored training parameters among devices. Explanation
- Partial fine-tuning. Updating selected existing parameters while freezing the others; the ASR starter freezes its acoustic tower and does not use LoRA. Explanation
- Patch merger. Combination of neighboring features into fewer input units. Explanation
- Pipeline parallelism. Distributing successive layer groups across stages. Explanation
- Polygon simplification. Reducing boundary vertices while bounding an allowed geometric change. Explanation
- Pooling. Combining several positions into a smaller summary representation. Explanation
- Precision. The fraction of predicted positive cases that are actually positive. Explanation
- Prediction contract. A specification of input fields, target meaning, output format, timing, and failure handling. Explanation
- Prefill. Processing the supplied prefix or prompt. Explanation
- Pretraining. Learning an initial representation or generative model from a broad objective. Explanation
- Pretraining from scratch. Training a newly initialized model rather than adapting pretrained weights. Explanation
- Probability calibration. Agreement between stated probabilities and observed outcome frequencies. Explanation
- Projector. Map from encoder features to a backbone's feature width. Explanation
- QAT. Training that accounts for quantization effects. Explanation
- QLoRA. Training floating-point low-rank adapters through a frozen quantized language-model base. Explanation
- Quantization. Representing selected numeric values with fewer bits under a specific runtime format. Explanation
- Query instruction. A task description added to a retrieval query in the format expected by an instruction-aware embedding model. Explanation
- Qwen3 embedding representation. The selected Qwen encoder's last-token-pooled, normalized vector with model-specific query formatting. Explanation
- R-squared. A regression comparison against predicting the mean, whose interpretation depends on the evaluation data. Explanation
- RAG. Retrieval-augmented generation: supplying retrieved external evidence to a model at inference time. Explanation
- Recall. The fraction of actual positive cases that the system finds. Explanation
- Recall@k. A retrieval measure of relevant results found among the first k candidates under the task's relevance definition. Explanation
- Recurrent state. Persistent summary updated as tokens arrive. Explanation
- Reference policy. A fixed baseline distribution used to measure how the trained policy changes completion likelihoods. Explanation
- Regression. Predicting a numeric quantity rather than a discrete class. Explanation
- Regularization. Constraints or training choices intended to improve generalization rather than only training fit. Explanation
- Reliability. In calibration, the relationship between predicted probability bins and observed outcomes. Explanation
- Reranker. A second-stage model scoring a smaller candidate set more precisely. Explanation
- Resampling. Changing an audio sample grid with suitable filtering; downsampling must suppress frequencies above the new Nyquist limit. Explanation
- Resident parameters. Complete weights that must be stored somewhere. Explanation
- Residual connection. Adding an earlier representation to a learned transformation. Explanation
- RMSE. Root mean squared error, emphasizing larger errors while retaining target units. Explanation
- ROC-AUC. A score-ranking measure comparing positive and negative examples across thresholds. Explanation
- RoPE. Rotary positional embeddings: positional relationships represented by rotations in attention coordinates. Explanation
- Router. Learned function that selects experts for token representations. Explanation
- Run manifest. A saved record of model/data identity, configuration, versions and execution details for one experiment. Explanation
- RVQ. Several codebooks used successively for discrete signal representation. Explanation
- Sample rate. The number of waveform amplitude measurements per second; changing a file header is not resampling. Explanation
- Seed. An initial state for pseudo-random generation. Explanation
- Semantic segmentation. Assigning a class to each pixel without necessarily separating individual instances. Explanation
- Sequence distillation. Using teacher-produced complete output sequences as student training targets. Explanation
- SFT. Supervised fine-tuning on examples of desired output conditioned on the supplied input. Explanation
- Shared expert. Expert used for all tokens alongside selected experts. Explanation
- Short convolution. A sequence-mixing operation over a local window, used with input-dependent gates in LFM2. Explanation
- Skip connection. A route carrying earlier features to a later stage, often preserving detail. Explanation
- Sparse attention. Attention over a selected subset of positions. Explanation
- Speaker diarization. Estimating which speaker spoke when; a separate task from recognizing the spoken words. Explanation
- Speculative decoding. Draft proposals checked by a target model. Explanation
- SSM. Structured state-space sequence model. Explanation
- Structured output. An output governed by explicit fields, types, labels, or other schema constraints. Explanation
- SWA. Direct attention limited to a moving local window. Explanation
- Target. The desired output used to define correctness during training. Explanation
- Target modules. The specific network modules selected to receive trainable adapters. Explanation
- Target scaling. Transforming target values during training and reversing the transformation for output. Explanation
- Teacher forcing. Training predictions on a supplied correct history, rather than the model’s own earlier generated mistakes. Explanation
- Temperature in calibration. A held-out-fitted logit scale used to improve probability calibration. Explanation
- Temperature in contrastive learning. A scale controlling the sharpness of contrastive candidate probabilities. Explanation
- Temperature in generation. A logit rescaling used to change a sampled token distribution. Explanation
- Tensor. A numeric array with shape, dtype, and device. Explanation
- Tensor parallelism. Splitting matrix operations across devices. Explanation
- Test. A held-out evaluation role used after development choices are fixed. Explanation
- Test set. Held-out data used after development choices are fixed. Explanation
- TF-IDF. Term frequency-inverse document frequency, a sparse representation weighting local term counts against corpus prevalence. Explanation
- Threshold. A selected cutoff converting a score into a decision. Explanation
- Tied embeddings. Reusing one parameter matrix for token input embeddings and vocabulary output projection. Explanation
- TinyML. Machine learning designed around tight device memory, compute, power, and latency limits. Explanation
- Tokenization. Converting text into the discrete token identifiers expected by a particular model. Explanation
- Tokenizer. A mapping between text and model token identifiers. Explanation
- Tool observation. The result returned by the external executor and supplied as context to the model. Explanation
- Tool schema. A structured specification of a tool name, purpose, argument types and constraints. Explanation
- Top-k routing. Selecting k expert assignments for each token. Explanation
- Training from scratch. Optimizing randomly initialized parameters rather than adapting pretrained weights. Explanation
- Transfer learning. Reusing a pretrained representation for a new task. Explanation
- Triplet loss. An objective comparing an anchor, a positive match, and a negative example. Explanation
- Truncation. Discarding input or target content beyond a length limit. Explanation
- Validation. An evaluation role used to compare candidates and select configuration or checkpoints. Explanation
- Validation set. Held-out data used to guide experiment choices. Explanation
- Virtual environment. An isolated Python package directory tied to a chosen interpreter. Explanation
- Voice activity detection (VAD). Predicting which audio regions likely contain speech, with a tradeoff between non-speech suppression and accidentally deleting quiet speech. Explanation
- Weight decay. A parameter-shrinking term applied by the optimizer. Explanation
- Word error rate (WER). Substitutions plus deletions plus insertions, divided by the number of reference words under a specified normalization/tokenization policy. Explanation