Training Your
Own Models
Browse the book
Chapter 16 / 4020 min read

16 Project make related text easy to retrieve

An embedding encoder turns an input into a vector designed for a downstream similarity objective. Useful application patterns include semantic search, near-duplicate discovery, clustering, matching queries to catalog items, and a frozen encoder feeding a small task-specific classifier. A dense retriever can feed a more expensive reranker or supply evidence to a separate answer-generating system. Its output is a representation or ranked result, not a guarantee that a generated answer is correct.

Use embeddings when semantic relationships matter and exact matching alone is insufficient. Preserve deterministic filters for permissions, dates, product variants, and structured constraints. A vector search should never bypass access control just because a private document is similar. Compare lexical, dense, and combined approaches on your actual queries.

Open a fresh terminal at the unpacked companion root before the commands below. They explicitly enter examples/embeddings; do not resolve that path relative to a previous project directory.

Imagine a user asks, “How do I stop future renewals?” while the correct help article says “Cancel a subscription.” Exact keyword matching may miss the relationship. An embedding model turns each text into a vector so useful pairs can receive high similarity. It is a representation model; it does not need to generate an answer.

First run the lexical baseline. A neural model should earn its extra complexity. Start from a fresh environment; do not run an import check before installing its dependencies.

Fresh environment for the CPU retrieval baseline

The following commands assume Python 3.12 is installed and a Linux/macOS shell starts in the downloaded book directory:

cd examples/embeddings
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements-cpu.txt
python -c "import numpy, sklearn; print(numpy.__version__, sklearn.__version__)"
python make_retrieval_data.py --out data/retrieval
python lexical_baseline.py --data data/retrieval --out runs/lexical

On Windows PowerShell, create the environment with py -3.12 -m venv .venv and activate .venv\Scripts\Activate.ps1, then use the remaining Python commands. Respect local execution-policy controls. Existing output directories should be replaced with new experiment names, not combined with new artifacts. The setup paths in this chapter are relative to examples/embeddings after the first cd.

The generator creates 60 fictional help documents for ten fictional products and three query variants per document. Products are split six/two/two across training, validation, and test. Training contains 108 query/positive pairs, but only 36 distinct positive documents. The shared templates make this a wiring fixture, not a serious retrieval benchmark.

The CPU-tested TF-IDF baseline achieved validation Recall@5 of 1.0 and MRR@5 of about 0.821 on this fixture. That is a useful warning: a neural model is not automatically necessary, and the dataset is deliberately easy. Real improvements require independently collected queries, relevant-document judgments, and realistic distractors.

What the keyword baseline represents

TF-IDF means term frequency-inverse document frequency. It represents a document with weighted counts: terms repeated within that document gain weight, while terms common across the corpus are downweighted. The supplied lexical baseline includes single words and adjacent two-word sequences called bigrams. Its vectors are sparse because most vocabulary entries do not appear in a given passage. A query is transformed with the same fitted vocabulary and scored against document vectors. This is a strong inspectable baseline when exact terminology carries relevance; it does not require a neural network.

Replace the fixture with your own query to document task

A JSONL file stores one JSON object per line. Keep three files in your data directory. These first rows form one consistent example:

corpus.jsonl contains documents that can be searched:

{"doc_id": "aster-password", "text": "Aster: How to change the password. Open Account then Security and choose Change password.", "group": "Aster", "split": "train"}

queries.jsonl contains questions with relevance judgments for evaluation:

{"query_id": "aster-password-q0", "text": "How do I change the password in Aster?", "relevant_doc_ids": ["aster-password"], "split": "train"}

train_pairs.jsonl contains the positive pairs used for adaptation:

{"query": "How do I change the password in Aster?", "positive": "Aster: How to change the password. Open Account then Security and choose Change password.", "doc_id": "aster-password", "group": "Aster"}

doc_id is a unique stable identifier. query_id identifies a query. relevant_doc_ids is a list because several documents can correctly answer one query; populate all verified positives rather than assuming only one. The training row's positive text must exactly match its referenced corpus document. Do not silently change whitespace, document content, or IDs in only one file.

group is the independent source unit held together across the split: the fixture uses product families; a real task might use source-document families or time-separated collections. Here every corpus group belongs entirely to train, val, or test; query judgments stay with the corresponding document split. All paraphrases of a query stay together. The shown row is training data, so add different groups labeled val and test with their own evaluation queries. Only train documents may appear in train_pairs.jsonl.

The search index contains the entire permitted corpus, including held-out documents, because these are the inference-time database. Their held-out query labels are not used for weight updates. Do not put sensitive or inaccessible documents into a shared index without appropriate access controls. For a different deployment question, such as future unseen corpora, define that stronger split deliberately rather than pretending this fixed-corpus exercise tests it.

For this row, the default neural query input is literally:

Instruct: Given a support question, retrieve the help article that answers it for the correct product.
Query:How do I change the password in Aster?

Its positive document is unprefixed:

Aster: How to change the password. Open Account then Security and choose Change password.

The instruction consumes part of the 128-token limit. The program adds it exactly once in training, validation, final evaluation, and live search. The corpus text itself is not rewritten with this query instruction.

What an embedding represents

A text encoder tokenizes the text, produces token-level representations, and pools them into one fixed-length vector. Mean pooling averages non-padding token representations; including padding would make the result depend on how it was batched. Other models use a designated token, last-token pooling, or more elaborate representations. Use the pooling the model was trained for.

The primary model is Qwen3-Embedding-0.6B, a compact member of the instruction-aware Qwen3 embedding family. It is public/non-gated at verification and marked Apache-2.0. It offers 1024-dimensional embeddings and an advertised 32K context, but this lab deliberately uses only 128 tokens. Model capacity is not a guarantee that a full 32K training batch fits a 24 GB card. M30

Its released module configuration uses a Transformer, last-nonpadding-token pooling, and normalization. Mean pooling is a useful concept to know, but it is not the recipe for this checkpoint. The shared query format is an instruction followed by Query: and the actual question; documents receive no instruction. The saved tokenizer handles the model's native EOS behavior. Do not apply a chat template or append a second EOS manually. M31 M32

The architecture configuration specifies 28 layers, hidden width 1024, 16 query heads, 8 key/value heads, head dimension 128, and feed-forward width 3072. Counting the published embedding/backbone tensor shapes implies approximately 595,776,512 parameters; this is a source/config-derived count, not an author-executed model load. The script prints the actual loaded parameter count for verification. It adapts the encoder through Sentence Transformers rather than constructing a text-generation head. M33

A short historical contrast: an older MiniLM mean-pooled encoder is much smaller and illustrates a different pooling convention. It is not the primary training route in this project. Model size, pooling, and query formatting must be chosen together rather than swapping names in an otherwise unchanged pipeline.

A tokenizer converts text into model-specific token IDs. Tokens are not necessarily words. A new tokenizer changes the meaning of the model's embedding lookup indices; you cannot replace one casually while keeping all pretrained weights. Keep casing, special tokens, query/document prefixes, truncation, pooling, and normalization consistent between training, indexing, and search. Some other retrieval models require task-specific prefixes; this lab uses the Qwen instruction format consistently in training, validation, indexing, and live search.

L2 normalization rescales a nonzero vector to length one. For normalized vectors, cosine similarity equals dot product, and squared Euclidean distance is 2 − 2 × dot_product; their rankings agree. Without normalization, vector lengths can change dot-product rankings. Do not train with one scoring rule and index with another accidentally.

Install the neural runtime then run a smoke experiment

The pinned, source-reviewed API target is PyTorch 2.8.0, Sentence Transformers 5.1.1, and Transformers 4.57.1. These meet the model card's requirements; they are explicit reproducibility versions, not a claim to be the newest releases. The loss, SentenceTransformer loader, and Qwen3 implementation were inspected at their tags. No neural packages were installed, no weights downloaded, and no neural training run during authoring. M14 M26 M34

Remain in the activated environment. Choose exactly one PyTorch route. On Linux/Windows with an NVIDIA driver and GPU compatible with the official CUDA 12.8 wheel:

python -m pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cu128
python -m pip install -r requirements-neural.txt
python -m pip check
python -c "import torch, transformers, sentence_transformers; print(torch.__version__, transformers.__version__, sentence_transformers.__version__); print('CUDA:', torch.cuda.is_available()); print('BF16:', torch.cuda.is_available() and torch.cuda.is_bf16_supported())"

For CPU-only Linux/Windows, replace the first installation command with python -m pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu, then run the same requirements and checks. For macOS CPU use python -m pip install torch==2.8.0 instead. This recipe does not implement MPS. Changing --device cannot give a CPU wheel CUDA support. For a GPU unsupported by these pinned wheels, select an official compatible stack and revalidate; do not force an incompatible build. The official versioned install page lists the supported wheel routes. M35

After reviewing the model license, begin with one update:

python train_embeddings.py --data data/retrieval --out runs/qwen3-smoke \
  --device cuda --amp-bf16 --batch-size 2 --max-length 128 \
  --epochs 1 --max-steps 1

Confirm the smoke checkpoint reloads and returns document IDs before attempting the longer run:

python search.py runs/qwen3-smoke "How do I change the password in Aster?" \
  --device cuda --k 3

This is a wiring check on a training-style question, not a generalization score. Then use a new output directory for adaptation:

python train_embeddings.py --data data/retrieval --out runs/qwen3-adapted \
  --device cuda --amp-bf16 --batch-size 2 --max-length 128 \
  --epochs 3 --lr 0.00001 --weight-decay 0.01 --scale 20

The script pins Qwen/Qwen3-Embedding-0.6B to verified revision 97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3, keeps trust_remote_code=False, and rejects a nonempty run directory. The first model load downloads about 1.2 GB of weights plus tokenizer/configuration; a saved FP32 checkpoint is larger. Leave generous disk and host-RAM headroom. Pinning an artifact does not replace license and provenance review. M36

CPU training uses the same command with --device cpu and without --amp-bf16. Full 0.6B backpropagation is slow on CPU and needs substantial RAM. For a first CPU neural retrieval experiment, save and index the unmodified baseline without allocating an optimizer:

python train_embeddings.py --data data/retrieval --out runs/qwen3-baseline \
  --device cpu --eval-only --batch-size 2 --max-length 128

The adaptation run updates all encoder parameters, so it is full fine-tuning of a pretrained representation, not scratch language pretraining. There are 36 unique positive documents, hence 18 updates per epoch at batch two and 54 across three epochs. The fixture is sufficient to exercise the machinery, not to claim adaptation quality. Use independently judged real query/positive pairs before drawing product conclusions.

The program first scores the unmodified model on validation and saves it as an eligible candidate. Each epoch samples one query variant per positive document, trains, and measures validation retrieval. It selects a new checkpoint only when validation MRR@5 improves. selected_epoch=0 means the unmodified model won. The chosen model is then validated and used for corpus indexing. Smoke, baseline, and training commands do not score test queries; final test evaluation is a separate explicit command after all tuning is done. Optimizer and original-model references are released before loading the selected checkpoint, avoiding a needless second training-sized model allocation.

The embedding experiment controls

Control Default Meaning
model and revision pinned in embedding_contract.py Verified Qwen3-Embedding-0.6B artifact; changing architecture requires reviewing its whole input/output contract
--epochs 3 Passes over unique positive-document groups
--batch-size 2 Actual contrastive pairs per update and encoding batch size
--max-length 128 Token budget including instruction and native special tokens; longer inputs are truncated
--lr 0.00001 AdamW learning rate for full-encoder adaptation
--weight-decay 0.01 Decoupled weight penalty
--scale 20 Cosine similarity multiplier before contrastive cross entropy
--instruction support-retrieval sentence Task description prefixed to queries, never documents; persisted for inference
--seed 42 Sampling/random-state control
--device cpu CPU or CUDA encoder/training placement
--amp-bf16 off CUDA BF16 autocast on supported hardware; master weights and optimizer state remain FP32
--max-steps 0 No update cap; use 1 for a smoke run
--eval-only off Save/evaluate/index the frozen baseline without training
--data, --out data/retrieval, runs/qwen3-embedding Dataset and artifact locations
search --k 5 Number of ranked result IDs to return

The loop clips gradient norm to 1 and deliberately omits a scheduler for the introductory short run. It uses SDPA without requiring a FlashAttention install, non-reentrant gradient checkpointing, use_cache=False, and AdamW foreach=False. There is no decoder-generation KV cache needed for embedding training. Gradient checkpointing saves activation memory by recomputing parts of the forward pass. BF16's numerical range generally avoids the FP16 gradient-scaling requirement; this code does not use an FP16 mode. M34

Match embedding architecture and supervision to the task

A single-vector encoder compresses a whole input into one vector and is efficient to index. It may blur a long document's small but important detail; chunking or a late-interaction/multi-vector architecture can preserve more detail at extra memory and serving cost. A cross-encoder reads query and document together and suits reranking a short list, but does not produce independently reusable document vectors in the same way.

Training a text encoder from random initialization requires learning language structure as well as your retrieval relationship. Fine-tuning an existing small encoder can focus a limited set of high-quality judgments on the latter task. For adaptation, the most valuable data are representative queries with verified positives, realistic confusions, and examples of what should not match. A million mechanically generated easy pairs can be less useful than a much smaller carefully judged set. The fixture demonstrates the format and loop; it does not estimate the needed count for your domain.

Build learning curves over independently labeled query families and inspect truncation rates, false negatives, and relevant-document coverage. For a multimodal encoder, you need aligned examples spanning the modalities, not arbitrary text and image collections placed side by side.

Contrastive learning turns a batch into a classification problem

For a batch of B query/document pairs, compute B query vectors and B positive-document vectors. Construct a B × B similarity matrix. Row i's correct candidate is document i; the other documents are treated as negatives. Cross entropy rewards the diagonal. This is multiple-negatives ranking loss, a common contrastive training objective. M14

With cosine similarity s, its query-i loss is:

−log( exp(scale × s(query_i, positive_i))
      / sum_j exp(scale × s(query_i, document_j)) )

scale=20 corresponds to temperature 0.05. Larger scale makes the distribution sharper and changes gradient behavior. It is not an accuracy setting. The script uses the library's loss to handle numerically stable cross entropy rather than directly implementing the exponentials shown for explanation.

A batch of size one has no competing documents and no useful contrastive training signal. The script discards singleton final batches and requires at least two distinct positives. It samples one query per document per epoch so the same positive document does not appear twice in one batch. The standard Sentence Transformers no-duplicates sampler is another useful tool for applicable trainer workflows. M15

A triplet loss instead uses an anchor, a positive, and an explicit negative. A common distance form is max(0, distance(anchor,positive) − distance(anchor,negative) + margin). The margin asks for separation, not a universal semantic distance. Easy negatives that already satisfy the margin contribute nothing; mislabeled hard negatives push the representation in the wrong direction. Triplet methods are a distinct training choice, not a reason to add arbitrary integer labels to a pair-based loss. M23

False negatives hard negatives and leakage

An in-batch negative is an assumption, not a verified fact. Two different articles may both answer “How do I reset access?” If one is paired as the positive and the other appears in the batch, the loss may punish a genuinely useful match. Exact-text deduplication catches identical strings, not semantic duplicates.

For real data, group equivalent documents and queries; represent multiple relevant documents in evaluation; construct batches that avoid competing positives where feasible; or choose a loss that explicitly handles multiple positives. Do not equate “not labeled positive” with “known irrelevant.” Audit mined examples manually before scaling them.

A hard negative is a plausible but wrong candidate. For the help-desk case, an article for the wrong product or an outdated policy can be hard. Mine candidates from training queries using a baseline retriever, then verify their irrelevance. Never mine from held-out test queries and feed those pairs back into training. If a hard negative is actually relevant, additional training can make the system worse while the training loss improves.

Split by source document, customer question family, or other genuine group before creating paraphrases. A test query that is a light rewrite of a training query is weak evidence. The fixture separates product groups but shares wording templates; its evaluation limitation is intentional and documented.

Retrieval has a subtle distinction from classification: the search corpus can contain held-out documents because they are the database being searched at inference time. Indexing their text is not the same as using their held-out query/relevance labels to train. This lab evaluates the fixed-corpus setting. If the claim is generalization to a future changing corpus, hold out time periods or document families accordingly and test that stronger setting.

Batch size means something different here

A larger real contrastive batch supplies more candidate negatives to each query, which can improve the learning signal until noise or false negatives dominate. Gradient accumulation does not automatically enlarge the in-batch negative set. Four separate batches of 16 pairs with accumulated gradients still present 16 candidates per loss computation, not one joint 64-candidate classification problem.

Ordinary attention activation memory depends on batch size and token length; the dense similarity matrix also grows with B squared. Long padded outliers can waste considerable space. Limit or bucket token lengths deliberately, inspect actual truncation, and reduce per-step batch size or sequence length when memory is tight. Use a cached contrastive loss when its larger negative pool is justified: gradient caching trades additional computation for lower activation-memory pressure, but does not make all memory costs disappear. The original gradient-cache paper and current library loss documentation explain this tradeoff. M16

On a 24 GB GPU, the small encoder and short default batches are a conservative starting experiment, not a tested fit guarantee. Run one forward/backward/update step with representative maximum-length examples, record peak allocated and reserved memory, leave room for evaluation and indexing, and only then increase batch size. The provided loop keeps FP32 master weights and optimizer state while optionally autocasting eligible CUDA operations to BF16. Parameters, gradients, and two Adam moments alone account for roughly 8.88 GiB for the config-derived parameter count, before activations, workspaces, allocator reserve, and any process sharing the GPU. That is a planning calculation, not a measured fit. The script records both peak allocated and reserved CUDA memory; the smoke run must establish whether your actual configuration fits. If it does not, reduce length/batch while retaining at least two contrastive pairs, or select a separately validated parameter-efficient method.

Evaluate the retrieval product not just its loss

The neural evaluation computes rankings against all 60 documents and reports:

  • Recall@k: fraction of the query's known relevant documents retrieved in the top k
  • MRR@k: reciprocal rank of the first relevant result, or zero if none occurs within k, averaged across queries
  • nDCG@k: discounted relevance accumulated near the top, divided by the ideal ordering; the example uses binary relevance, while real datasets can use graded judgments

The implementation supports multiple relevant document IDs and deterministic tie ordering. It does not assume the index contains only one positive and a few handpicked negatives. Sentence Transformers provides information-retrieval evaluators for larger workflows and additional metrics. M18

A reranker jointly reads a query and each retrieved candidate, often scoring relevance more accurately at greater latency than independently encoded vectors. A useful pipeline is lexical and/or dense retrieval → candidate merge → reranking → final selection. The reranker cannot recover an answer that the first-stage candidate set omitted. Measure first-stage recall, final ranking quality, and end-to-end latency separately.

Similarity is not a calibrated probability of correctness. A cosine score of 0.8 is not automatically an 80% chance of answering the question. Validate any “no answer” threshold using real answerable and unanswerable queries. Include product names, identifiers, negation, numbers, dates, languages, and near-duplicate policies in the error set. Dense similarity alone can be weak on exact codes; lexical retrieval remains a valuable complement.

Use the trained index consistently

The output directory contains:

  • best/: the selected encoder, tokenizer, last-token pooling configuration, and associated model files
  • index_contract.json: exact instruction, preprocessing, dimensions, and SHA-256 identities for model and index files
  • corpus_embeddings.npy: normalized vectors produced by that selected model
  • corpus_ids.json: the row-to-document mapping
  • validation_rankings.json: inspectable validation result IDs; final test artifacts are created only by the explicit test command
  • metrics.json: configuration, validation history, selected validation scores, input data hashes, and available memory measurements

The short trainer saves inference-ready selected weights, not optimizer/scheduler state or a resumable training position. Reloading best/ for search is not resuming interrupted training. A resumed training experiment would need an explicitly implemented optimizer/random-state checkpoint strategy and its own checks.

Search it with:

python search.py runs/qwen3-adapted \
  "In Iris, I need to stop future renewals." --device cpu --k 5
python search.py runs/qwen3-adapted \
  "In Iris, I need to stop future renewals." --device cuda --k 5

Use runs/qwen3-baseline instead if you created the frozen baseline. The first command encodes on CPU; the second encodes on CUDA. Both load FP32 weights, independently of the training autocast choice. Each loads the selected local model directory, including tokenizer and last-token pooling, with the saved sequence-length limit, left padding, and exact training instruction. The query prefix is applied once; documents remain unprefixed. SHA-256 checks reject mismatched weights/tokenizer/configuration/index files, adding startup I/O in exchange for consistency. The document index stays in NumPy on CPU in this small demo; only the encoder changes device. Vectors are normalized on both paths, and the output contains document IDs and similarity scores. These neural commands were syntax-checked, not executed. For a larger GPU index you would need a separately selected indexing backend and its own correctness checks.

When you change the encoder, pooling, normalization, tokenization, or document chunking, rebuild the index or implement a carefully tested migration. Old document vectors and new query vectors are generally incompatible. Store model revision and preprocessing identity alongside every index version. Protect sensitive documents and query logs; embeddings are derived data, not an automatic anonymization method.

The lab uses an exact NumPy dot-product search. For a large corpus, measure approximate-nearest-neighbor index recall against exact search on a sample. Quantized vectors and compressed indexes reduce memory but can change ranking. Approximation errors compound with encoder and reranker errors, so evaluate the complete deployed path.

Reserve the test set for one final comparison

During development, both the lexical command and neural training commands report validation only. The one-update smoke run and frozen-baseline run must not become early peeks at the held-out test. When the model, instructions, chunking, index settings, and thresholds are fixed, run:

python evaluate_saved.py runs/qwen3-adapted --data data/retrieval \
  --split test --device cuda --batch-size 2
python lexical_baseline.py --data data/retrieval \
  --out runs/lexical-final --split test

CPU final evaluation uses --device cpu; for a deliberately selected frozen encoder, substitute its run directory. The neural evaluator checks corpus identity against the stored index contract and writes test_metrics.json and test_rankings.json. Compare the fixed systems, report uncertainty and representative failures, and do not tune on these labels afterward. evaluate_saved.py defaults to --split val when the flag is omitted. A serious iteration after examining test failures needs a fresh final holdout.

Extend embeddings to images and multiple modalities

The same pattern applies beyond text. An image encoder maps photographs to vectors for visual similarity or a small downstream classifier. The definition of “similar” matters: identical products, matching texture, same species, and the same individual object are different objectives. Image augmentations define which variations the representation should ignore; an aggressive crop can remove the very identity signal you wanted to preserve.

A dual-encoder multimodal system has, for example, an image encoder and a text encoder that map into a shared dimensional space. Contrastive training encourages paired images and descriptions to align. CLIP is a primary example of this training pattern. It is not achieved by comparing arbitrary image vectors with unrelated text vectors that happen to have the same length. M17

For a small custom multimodal task, begin with an appropriate pretrained aligned model and frozen-feature baselines. Fine-tune on rights-cleared matched pairs only if validation shows a need. Captions that omit the decisive visual detail are weak supervision for that detail. Multiple images can match a caption, creating the same false-negative issue seen in text retrieval. Evaluate image→text and text→image separately if both directions matter.

Do not use a retrieval embedding as a substitute for a pixel mask. A single pooled image vector intentionally compresses spatial information; precise boundaries require an output architecture and labels that preserve it.