26 Project teach an image generator a small visual style
An image diffusion system generates a new image by iteratively transforming a noise-like tensor, usually guided by text or other inputs. It is useful when you want plausible visual alternatives: concept art, original asset variations, product-scene mockups, or controlled edits. It is unsuitable when invented details would be mistaken for evidence, such as filling in a supposedly factual medical or forensic image. For counting defects, recognizing objects or assigning a class to each pixel, a classifier, detector or segmentation model is usually the more direct design.
Common application patterns are:
- Text → candidate images → human selection. Inputs include prompt, dimensions, seed and generation settings. Outputs include image files and a generation record. Add checks for prohibited content, personal likenesses and downstream use; generating an image does not establish that it is accurate or rights-cleared.
- Reference image → bounded variation. An image-to-image pipeline encodes a reference and introduces a chosen amount of noise before generation. Stronger modification can erase identity or layout. This is an inference pipeline choice, not proof that you trained a new subject.
- Image plus aligned mask → localized edit. An inpainting-capable pipeline receives both. Specify which region may change, preserve the unedited source, and inspect edges and supposedly preserved regions afterward. The ordinary text-to-image LoRA trainer below does not learn this interface automatically.
- Layout/depth/edge condition → controlled generation. A purpose-built conditioning model can guide structure. Keep the control map's coordinate system, resolution and semantics aligned with the image. Text prompts alone do not enforce exact geometry.
A deployment contract should name the base revision, adapter revision, input types, supported sizes, expected latency, output format and limitations. Keep generation asynchronous when it takes longer than an interactive response. Save candidates with their settings, allow users to reject them, and separate a model's technical success from approval for publication. The first project covers the simplest pattern, text to image, with an adapter for a narrow visual style.
What you will make. A small adapter that nudges an existing image generator toward an original paper-card illustration style. You will create the training images locally, make a baseline image, train only a small fraction of the generator's weights, and compare the result with the unchanged model. Nothing is uploaded by the commands in this project.
Why start here? You can see whether the experiment works. The base generator already knows many objects, so you do not need to teach it the entire visual world. Later, the audio project removes this shortcut and builds a very small generator from random weights.
What is and is not verified. The companion data generators and data checks were executed on CPU; Python syntax and shell syntax were checked. The image training command was checked against the actual Diffusers v0.36.0 source. No GPU training or image generation was executed for this book. The software pins are a deliberately conservative candidate environment, not a claimed end-to-end tested lockfile. The single-process trainer supports a conservative one-GPU setup; a 24 GB RTX 3090/4090 is a candidate target for this Sana LoRA configuration, but you must measure the first run. The model choices were checked against primary sources on 2 October 2026.
Image project step 1 create something the model can learn
All commands in this project run from the extracted
examples/diffusion folder. Set an absolute path once,
replacing the placeholder with your actual extraction location, and
return to it before beginning:
export DIFFUSION_EXAMPLES="/absolute/path/to/extracted/examples/diffusion"
cd "$DIFFUSION_EXAMPLES"
Do not paste the placeholder unchanged. If another project's virtual
environment is active, run deactivate first. Use Linux or
WSL2 with a working NVIDIA driver. A full CUDA toolkit is not normally
required merely to run an official PyTorch CUDA wheel; custom compiled
extensions can add that requirement. Do not install a random CUDA
package to fix a driver problem.
Make a virtual environment first. The base model is several gigabytes, and caches, checkpoints and a second environment add up: keep at least 50 GB of free SSD space for this chapter's experiments. That is a planning allowance, not the measured download size of a particular snapshot.
python3.11 -m venv .venv-image
source .venv-image/bin/activate
python -m pip install --upgrade pip
python -m pip install torch==2.7.1 torchvision==0.22.1 \
--index-url https://download.pytorch.org/whl/cu126
python -m pip install -r requirements-image.txt
python -m pip check
python -c 'import torch,diffusers,peft,transformers; print(torch.__version__, torch.cuda.is_available()); print(diffusers.__version__,peft.__version__,transformers.__version__)'
The exact torchvision/PyTorch pair follows PyTorch's versioned
installation instructions D27. A newer driver
can usually run a wheel built for an older supported CUDA runtime; check
your driver rather than equating the number printed by
nvidia-smi with an installed toolkit. The example targets
Ampere/Ada cards, not every GPU sold with 24 GB.
Pin the trainer and the library together. The
companion requirements use Diffusers 0.36.0, PEFT 0.17.1, Transformers
4.55.4, Accelerate 1.10.1 and Hugging Face Hub 0.34.4, plus the
tokenizer's SentencePiece dependency. These satisfy the inspected
library's declared minimums, but the complete environment still needs
the import and smoke tests on your machine D34. Save the resolved environment with
python -m pip freeze > image-environment.txt after it
works. This version is a reproducible recipe target, not a claim that it
is the newest release.
Sana's frozen text encoder and autoencoder are additional models, even though we train only adapters in the 1.648B-parameter image transformer. The complete pipeline exceeds three billion parameters. Budget system RAM as well as VRAM: 32 GB is a reasonable starting planning allowance and 64 GB offers more headroom for model loading and offloading, but neither is a measured requirement of this recipe. The trainer also prepares the processed image tensors in host memory. We print actual component parameter counts during the baseline load rather than labeling the whole pipeline “1.6B.”
Create 240 small, original illustrations:
python make_image_cards.py --out data/papercards --per-object 60
python check_image_data.py data/papercards
The generator draws houses, mugs, flowers and balloons with a fixed paper-like treatment. It creates 192 training, 24 validation and 24 test images. Each file is rendered natively at 1,024 × 1,024 pixels, with geometry scaled to the larger canvas rather than upsampling a small picture. This is intentionally a limited exercise: success on four procedural object families does not demonstrate that you have trained a broad commercial illustration model. Open a contact sheet or at least several files from each split before training. The correct first test is whether you like, understand and can legally use the data.
The layout is:
data/papercards/
provenance.json
train/
house_0000.png
...
metadata.jsonl
val/
...
metadata.jsonl
test/
...
metadata.jsonl
One line of metadata.jsonl looks like this:
{"file_name":"house_0000.png","text":"a blue house, rivetpaper style, centered on cream paper","group":"house-0"}
file_name connects a picture to its caption.
text says what is visible and names the desired style.
group keeps related originals out of different splits. The
upstream imagefolder format reads the image-caption connection; the
training script ignores our extra grouping field D16. Keep the validation and test folders outside
the training folder, and pass only data/papercards/train to
the trainer. Pointing the trainer at the whole tree can change split
discovery and risks training on unintended examples. Make the selected
training directory explicit.
Image data captions groups crops and masks
Replace the generated pictures with your own art or product photos after the pipeline works. A small style experiment might start with 50-200 genuinely varied originals, but that is a project-size suggestion, not a statistical minimum. Two hundred near-identical frames do not provide two hundred independent examples.
For a style adapter, vary object, background, layout and color. Otherwise the model can learn “this style means a centered red mug.” Describe visible differences in captions. For a subject adapter, capture the same subject from different views, distances and backgrounds so that identity and scenery are not inseparable. Split a photographic session or video into one partition as a group. Burst shots, crops of the same original, and lightly edited duplicates belong together.
A trigger such as rivetpaper style is a consistent
phrase, not a newly invented tokenizer token. It will be broken into
existing tokens. The adapter learns associations involving those tokens;
it does not magically add a dictionary entry. If every training caption
includes the trigger, test both with and without it to see how much the
style spreads into unrelated prompts.
Resizing and cropping are changes to the example, not neutral housekeeping. The beginner command resizes the shorter edge and center-crops a square. A wide image of two people may lose one person; a caption that still says “two people” is then wrong. Inspect the processed crop. Do not flip images with lettering, asymmetric logos, medical laterality or other direction-sensitive meaning just because a tutorial enables random flips.
An aspect-ratio bucket groups similarly shaped
images into batches, such as portrait, square and landscape. It reduces
destructive cropping without padding every image to the largest shape.
The upstream script used here does not implement arbitrary aspect-ratio
bucketing; adding a --buckets flag will not make it do so.
Move to a trainer with explicit, verified bucket support only after the
square experiment is understood. Hold total pixel area approximately
constant when comparing bucket shapes because changing shape can also
change memory requirements.
A mask is an aligned map indicating a region, often one value for editable pixels and another for preserved pixels. It is needed for an inpainting-specific training objective, or for weighting a loss to a region. Merely placing a mask file beside an image does not activate masked training in this LoRA script. Edge maps, depth maps, segmentation maps and poses are other forms of conditioning. They require the corresponding model inputs and training procedure. A segmentation model that labels pixels is solving a different task from a diffusion model that generates pictures; use the discriminative-model project when labels, rather than invented pixels, are your goal.
Image project step 2 download a fixed base and make a baseline
Use the official
Efficient-Large-Model/Sana_1600M_1024px_BF16_diffusers
checkpoint. Sana is a maintained efficient image-transformer family, and
the official Diffusers adaptation example specifically targets this BF16
checkpoint. Its card lists a 1.648B image transformer, a frozen
Gemma2-2B-IT text encoder and a 32× spatial-compression DC-AE
autoencoder. The Sana weights have Apache 2.0 terms, the Gemma component
has additional terms, and the card states research-oriented intended
use. Read the actual current terms and restrictions before use D30 D35 D36.
The official example establishes a real LoRA training route. It does not establish a measured 24 GB peak for our exact custom-caption dataset. Our single-process, batch-1, checkpointed/offloaded configuration is an unexecuted candidate designed to be measured before scaling. We chose the 1.6B BF16 variant because it is the dedicated checkpoint in that training recipe; a smaller 590M Sana variant exists, but substituting it also changes precision and model configuration and should be a separate experiment.
Download a revision-pinned local snapshot. This code resolves the
current model revision once, records it, and downloads that exact
revision. The recorded SHA is what reproduces the selection later. It
does not promise that every future run resolves the same
main revision.
python - <<'PY'
import json
from pathlib import Path
from huggingface_hub import HfApi, snapshot_download
repo = 'Efficient-Large-Model/Sana_1600M_1024px_BF16_diffusers'
sha = HfApi().model_info(repo).sha
path = snapshot_download(
repo_id=repo, revision=sha,
allow_patterns=['model_index.json', 'transformer/*', 'vae/*',
'text_encoder/*', 'tokenizer/*', 'scheduler/*'],
ignore_patterns=['*.onnx', '*.msgpack', '*.h5'],
)
Path('base_path.txt').write_text(path)
Path('base_revision.json').write_text(json.dumps({'repo':repo,'revision':sha},indent=2))
print(path)
PY
export BASE="$(cat base_path.txt)"
python sample_image_lora.py --base "$BASE" \
--prompt 'a blue house, rivetpaper style, centered on cream paper' \
--seed 17 --out samples/baseline-house.png
The download filter avoids unrelated root-level checkpoints, but
component folders may contain multiple precision variants and consume
substantial disk space. Our loader explicitly requests the BF16 variant;
do not delete weight shards merely because their names look similar.
Keep the model's configuration, scheduler, tokenizer, text encoder,
autoencoder and denoiser together; a lone .safetensors file
is not necessarily a complete pipeline.
A seed initializes a pseudorandom number generator. Keeping it fixed makes a before/after comparison more informative because the starting noise is the same. It does not guarantee bit-for-bit output across different GPU architectures, library versions or kernels. Record the software, model revision, prompt, seed, image size, sampler and inference-step count.
Before training, also generate a no-trigger prompt and one unfamiliar object, such as “a bicycle, rivetpaper style, centered on cream paper.” Keep these baseline files. If the base already does what you need, an adapter may be unnecessary.
Image project step 3 one small training run
Get the training script from the same version as the installed library:
git clone --branch v0.36.0 --depth 1 https://github.com/huggingface/diffusers.git diffusers-src
git -C diffusers-src rev-parse HEAD > diffusers-source-commit.txt
python patch_image_trainer.py diffusers-src
git -C diffusers-src diff > image-trainer.patch
export DIFFUSERS_SRC="$PWD/diffusers-src"
export DATA="$PWD/data/papercards"
export OUT="$PWD/runs/papercard-smoke"
STEPS=20 WARMUP=5 bash train_image_lora.sh
The small guarded patch skips the upstream trainer's final allocation of another full FP32 pipeline when neither final validation nor upload was requested. The adapter is already saved at that point. We evaluate in a fresh process using the matching BF16 pipeline instead, avoiding an unnecessary host-memory peak and the upstream final reload's different transformer dtype. Save the patch with the upstream source SHA D31.
This smoke test checks loading, preprocessing, forward computation,
backward computation and saving. It is not long enough to establish
image quality. A loss value appearing on screen is not sufficient:
verify the run writes pytorch_lora_weights.safetensors and
that sample_image_lora.py can load it. The command uses
local TensorBoard logs and does not enable upload flags.
Run the actual reload check now, before spending time on the longer run:
python sample_image_lora.py --base "$BASE" --lora runs/papercard-smoke \
--prompt 'a blue house, rivetpaper style, centered on cream paper' \
--seed 17 --out samples/smoke-house.png
The smoke image is a file-loading/functionality check, not evidence that 20 steps learned the style. If either the adapter or its base manifest fails to load, fix that now.
Then start a separate deliberate run:
export OUT="$PWD/runs/papercard-lora"
STEPS=1000 WARMUP=100 bash train_image_lora.sh
The wrapper saves book_run_manifest.json, binding the
adapter to the exact base repository/revision, model configuration
hashes, tagged-and-patched trainer, wrapper settings, caption metadata
and training-image hashes. Resume rejects changes in those material
inputs; inference checks the adapter/base binding before loading. Keep
the base snapshot directory's revision name and
base_revision.json when moving the experiment. The large
base-weight files are treated as an immutable Hub snapshot rather than
rehashed on every image request; do not edit them in place.
The companion shell script expands to these important settings:
1,024-square images, per-device batch 1, accumulation 4, rank and alpha
8, learning rate 0.0001, BF16 mixed precision, gradient checkpointing
and frozen-component offloading, and a checkpoint every 250 optimizer
updates. All arguments were checked against the tagged Sana
implementation D31. We use
--dataset_name pointing to the local imagefolder and
--caption_column=text. The required
--instance_prompt is a fallback; our nonempty per-image
captions take precedence. Do not use --instance_data_dir
for this metadata-bearing folder: that branch treats files as images
rather than loading captions. The script is named DreamBooth, but also
accepts a captioned style dataset without enabling prior
preservation.
What these image training settings actually change
- Batch 1: one image is processed at a time on the GPU. This is the first memory lever.
- Accumulation 4: gradients from four successive microbatches are added before an optimizer update. With one GPU the effective batch is 4. It does not put four images in memory simultaneously and does not make a single oversized image fit.
- 1,000 optimizer steps: roughly 4,000 image presentations, or about 20.8 average passes over 192 examples. This is a starting experiment, not the correct number for every dataset. More training can make a model worse.
- Rank 8: each selected weight update is constrained to a low-dimensional correction. Higher rank gives more adjustable capacity and increases adapter memory; it does not add training examples.
- Learning rate 0.0001: the scale of parameter changes. A rate that works for LoRA is not automatically appropriate for updating every base weight.
- Warmup 100: gradually reaches the selected learning rate during the first 100 optimizer steps. The learning-rate schedule controls optimizer updates. It is unrelated to the diffusion noise schedule.
- BF16 mixed precision: uses bfloat16 for the image transformer and Gemma text encoder in this recipe; the autoencoder stays float32. These component-specific choices follow the inspected implementation. BF16 has a different numeric range/precision trade-off from FP16. A filename variant and a runtime arithmetic dtype are related but separate settings.
- Gradient checkpointing: discards selected intermediate activations and recomputes them during backward. It exchanges more computation for less activation memory.
- Offload: moves the frozen text encoder and autoencoder back to CPU when their turn is finished. This reduces simultaneous GPU residency and adds host RAM use and CPU-GPU transfers. It is not the same as training the text encoder or making the network smaller.
- 128 prompt tokens, no complex-human instruction: use the same limit and prompt preparation in training, baseline inference and adapted inference. The pipeline has an instruction-prefix default that this recipe deliberately disables; inconsistent preprocessing would confound the comparison.
- Gradient clipping 1: bounds the gradient norm before updating, limiting unusually large updates. It does not repair an invalid objective or corrupted data.
- Checkpoint every 250: saves resumable state. The final adapter is a deployment artifact; resumable checkpoints additionally preserve training state. Do not assume those files are interchangeable.
With this script, resuming an interrupted run uses the same output
folder and RESUME=latest. Keep the original configuration
and dataset unchanged. Example:
RESUME=latest STEPS=1000 WARMUP=100 bash train_image_lora.sh
Resume counts toward the total target of 1,000 steps, not 1,000
additional steps. This upstream trainer resumes optimizer/model state
but begins iterating from the stored epoch rather than exactly skipping
to the next old minibatch. Treat it as recovery, not bit-for-bit replay
of an uninterrupted data order. A finished checkpoint may be useful even
if you ultimately choose an earlier one by evaluation. At this inspected
tag, checkpoint folders also contain a saved LoRA file, so you can pass
--lora runs/papercard-lora/checkpoint-250 to the inference
script for that candidate.
Just enough diffusion what target did the network learn
In next-token language modeling, a model predicts a distribution over next token IDs. The Sana denoiser predicts a continuous tensor that describes movement along a path between a clean image latent and noise. The caption guides the prediction but is not the output target.
Let z0 be a clean compressed image and
epsilon be freshly sampled Gaussian noise of the same
shape. Choose a schedule level sigma and mix them:
z_sigma = (1 - sigma) * z0 + sigma * epsilon
training target = epsilon - z0
This is the flow-matching convention used by the inspected Sana trainer. At sigma zero the mixture is clean; at sigma one it is noise. The target is the derivative of that straight interpolation in the clean-to-noise direction. Generation follows the learned field in the reverse direction, using a compatible sampler. The network is not trained to predict a word or simply to emit the clean image in one step. Do not substitute a DDPM noise-prediction loss without also changing the model/objective contract D03 D31.
The trainer solves one randomly chosen path-level problem per example, using mean squared error against this movement target. Fresh noise and path levels create useful corruption variety, not new independent clean images. Generation instead begins with random noise and applies repeated model/solver updates. A thousand optimizer updates, a training schedule's available time levels and twenty inference steps are three different quantities.
The latent autoencoder and text encoder each have one job
Sana does not run its image transformer directly on every 1,024 ×
1,024 × 3 color value. The pretrained DC-AE autoencoder
compresses the image spatially by 32, and a decoder maps generated
latents back to pixels. A 1,024-square image therefore has a 32 × 32
latent spatial grid; channel count and scale come from the downloaded
model configuration. DC-AE is not interchangeable with Stable
Diffusion's older variational autoencoder. The Diffusers pipeline calls
its autoencoder component vae, but that attribute name does
not make every architecture a variational autoencoder D30 D32 D33.
The frozen Gemma text encoder converts the caption
into contextual vectors and an attention mask. The denoiser uses these
as a condition. It does not generate an enhanced caption in this recipe:
both training and inference set
complex_human_instruction=None, use the same tokenizer and
cap the encoded caption at 128 tokens. Token IDs, hidden-state width,
attention mask, negative-prompt convention and text normalization all
belong to the conditioning contract.
Caching can remove repeated work when representations are keyed to
their source example. We deliberately do not set
--cache_latents here: the inspected trainer caches by batch
position while its loader shuffles. With distinct captions, that can
pair a cached image latent with another image's caption on a later pass.
Keeping the cache disabled preserves alignment without pretending the
optimization is free. Any future cache should key image hash, caption,
crop/augmentation, encoder revision, latent scaling and precision
together D31.
The safe placement pattern in this script is: bring the text encoder to GPU for the current caption, move it back to CPU, bring the autoencoder to GPU for the current image, move it back, then perform the transformer update. Gradient checkpointing handles part of the remaining activation memory. This is slower than well-designed persistent caches, but simpler to inspect. The full models still exist in host or device memory; offloading does not erase their parameter counts.
A useful representation test is an autoencoder round trip: encode an input, decode it, and inspect the reconstruction. If fine lettering or texture is lost in that representation, training the denoiser longer cannot recover information the latent does not retain. Improving or replacing the autoencoder is a separate project.
LoRA is an update method DreamBooth is a personalization method
A full fine-tune allows many or all denoiser weights to move.
LoRA freezes the original weight matrix W
and learns two thin matrices, often written A and
B, whose product forms the correction:
W_effective = W + scale * (B @ A)
For a square 1,024 × 1,024 weight, full training has 1,048,576
adjustable entries. Rank 8 uses
8 × 1024 + 1024 × 8 = 16,384 entries for the correction,
before any additional adapter details. That is a simple parameter-count
illustration, not total model memory. You still store the frozen network
and run it to obtain gradients for the adapter. The official LoRA guide
explains the supported Diffusers integration D08.
DreamBooth is a strategy for associating a small collection of subject images with a prompt identifier, commonly including a class-based prior-preservation objective. It can be combined with LoRA or with fuller weight updates. Thus “LoRA or DreamBooth?” is not an either/or distinction. For ten images of one personal object, inspect a DreamBooth-specific recipe; for the captioned style dataset here, we use per-image captions and leave prior preservation disabled. Prior preservation helps resist replacing a broad class with one subject, but it does not guarantee against memorization D07.
Training the autoencoder, text encoder and denoiser together from scratch would ask the same small dataset to supply several kinds of knowledge. It greatly expands memory, data and debugging demands. A tiny 32- or 64-pixel pixel-space diffusion model can be an excellent mechanics experiment on one GPU; a general text-to-image foundation model from scratch is not a realistic beginner target for one 24 GB card.
Image project step 4 evaluate adaptation instead of cherry picking
python sample_image_lora.py --base "$BASE" --lora runs/papercard-lora \
--prompt 'a blue house, rivetpaper style, centered on cream paper' \
--seed 17 --out samples/adapted-house.png
python sample_image_lora.py --base "$BASE" --lora runs/papercard-lora \
--prompt 'a bicycle, rivetpaper style, centered on cream paper' \
--seed 17 --out samples/adapted-bicycle.png
Image inference on GPU and CPU
The inference script reconstructs the complete pipeline from the same local base snapshot and then loads the trained adapter. It uses the snapshot's tokenizer, text encoder, latent scaling, decoder and inference scheduler configuration. The training code constructs a flow-matching noise/path scheduler, while the pipeline can use a different compatible numerical solver at inference; matching the flow-prediction semantics is essential, not forcing identical scheduler class names. Text-to-image generation needs no input photograph; it starts from new noise, with the prompt encoded in the same vocabulary used during adaptation. The DC-AE decodes the final latent; the pipeline converts the image values to a PIL image and the script saves PNG. The adjacent JSON records prompt, device, dtype, seed, steps and scheduler configuration.
GPU inference is the default above. The companion implementation offers CPU execution through ordinary PyTorch operations, with all components in float32. GPU execution uses BF16 for the transformer/text encoder, FP32 for DC-AE and model CPU offload. CPU uses no CUDA offload hooks and can need substantially more system RAM; it is not equivalent in speed or memory to the GPU path:
python sample_image_lora.py --base "$BASE" --lora runs/papercard-lora \
--prompt 'a blue house, rivetpaper style, centered on cream paper' \
--seed 17 --device cpu --steps 20 --out samples/cpu-house.png
The CPU command is a float32 implementation path, not a measured
latency or universal-kernel-support promise. The official model examples
target CUDA. On CPU this multi-billion-parameter pipeline needs
substantial RAM and can be very slow; it is an optional diagnostic
fallback, not the recommended production path. A saved adapter does not
require optimizer state for inference, but it still requires its
compatible base. CPU and CUDA random generation and arithmetic may yield
different pixels even with the same seed. Compare semantic behavior, not
byte equality. Copy base_revision.json,
book_run_manifest.json, the environment record, adapter and
evaluation prompts together when moving the project to another
machine.
Make a fixed evaluation grid before deciding which checkpoint is best. Suggested rows are familiar objects with the style phrase, unfamiliar objects with the phrase, familiar objects without the phrase, and prompts that explicitly change the background, color or layout. Suggested columns are four fixed seeds. Generate the same grid with the base model and every serious checkpoint candidate.
Score separate questions rather than giving one vague score:
- Content: is the requested object present? Are requested colors and counts respected?
- Style: is the paper treatment learned beyond copying the training composition?
- Control: can the prompt alter background and layout? Does the trigger have an understandable effect?
- Diversity: do different seeds produce genuinely different acceptable designs?
- Retention: do ordinary prompts without the trigger still work?
- Memorization: does an output reproduce a particular training picture, including unusual placement or defects?
Use the held out folders in checkpoint selection
The following commands read the actual validation
metadata.jsonl, select eight evenly spaced records to cover
the generator's object families, and create two fixed-seed samples per
caption. The pipeline loads once per command. The helper writes a
cases.csv linking each generated image to its held-out
reference image, plus blank fields for your rubric scores. It does not
invent scores.
python evaluate_image_grid.py --base "$BASE" \
--metadata data/papercards/val/metadata.jsonl --split val \
--limit 8 --seeds 17,29 --out evaluations/base-val
python evaluate_image_grid.py --base "$BASE" --lora runs/papercard-lora/checkpoint-250 \
--metadata data/papercards/val/metadata.jsonl --split val \
--limit 8 --seeds 17,29 --out evaluations/step250-val
python evaluate_image_grid.py --base "$BASE" --lora runs/papercard-lora \
--metadata data/papercards/val/metadata.jsonl --split val \
--limit 8 --seeds 17,29 --out evaluations/step1000-val
Open each CSV's generated/reference paths and score content and style consistently. The reference picture is a style/content example for human review, not an input to the text-to-image model and not an exact pixel target. Compare diversity across the paired seeds and check suspicious outputs against the training set. The held-out drawings have new random visual variations, but their simple captions can repeat training captions. This validates narrow within-generator behavior, not new semantic categories. The separate bicycle/no-trigger/layout prompts probe broader transfer and control.
Choose the checkpoint using validation only. Then generate the same untouched test-caption cases for that selected checkpoint and the base. For example, if checkpoint 250 won:
python evaluate_image_grid.py --base "$BASE" \
--metadata data/papercards/test/metadata.jsonl --split test \
--limit 24 --seeds 17,29 --out evaluations/base-test
python evaluate_image_grid.py --base "$BASE" --lora runs/papercard-lora/checkpoint-250 \
--metadata data/papercards/test/metadata.jsonl --split test \
--limit 24 --seeds 17,29 --out evaluations/selected-test
Review those test reference/generated pairs once with the same rubric. Report the outcome even if it is worse. Selecting another checkpoint based on the test results turns that split into more validation data. Use fresh output directories for each evaluation; the helper refuses to mingle new outputs with an existing nonempty run.
Loss is useful for discovering broken training, but lower flow-target error does not automatically mean better pictures. For a tiny dataset, a large-sample distribution metric such as FID is unstable and can obscure the practical question. Start with reproducible side-by-side inspection, a clear rubric and nearest-neighbor review. Image-embedding similarity can help find suspicious training matches, but no threshold certifies originality.
For memorization checks, compare outputs to all training images, including crops and resized versions. Exact file hashes catch duplicates in the dataset but not visually near-identical images. If the model repeatedly reconstructs a training composition, stop earlier, reduce the learning rate or rank, remove repeated images, and add more independent variety. Do not call a memorized output “a successful style transfer” simply because it looks polished.
Guidance belongs mainly to sampling
Classifier-free guidance combines two predictions, conditioned and unconditioned. A common convention is:
prediction = unconditioned + g * (conditioned - unconditioned)
At g = 1, this expression is simply the conditioned
prediction; higher values extrapolate toward the condition. Training a
new model for this mechanism usually includes dropping the condition on
some examples, giving it an unconditional case to learn. The Sana
adaptation script uses a pretrained guided model; do not assume it has a
configurable caption-dropout flag. The audio script later implements
condition dropout explicitly D04.
A high guidance setting can increase apparent prompt adherence while damaging naturalness or reducing diversity. Try a small fixed set, such as 3, 5 and 7, while holding the checkpoint and seed fixed. That is an experiment, not a universal quality ranking. Do not change guidance between baseline and adapter comparisons unless that change is itself the experiment. Some distilled models use different guidance mechanisms or expected values; their cards and scheduler settings take priority over this generic formula.
Image troubleshooting and a 24 GB decision ladder
First measure the running process with nvidia-smi.
Record peak training use as well as separate sample-generation use. A
run can train successfully and fail while creating a large validation
batch. CUDA allocator reservations and process memory are related but
different measurements; record which you used.
| Symptom | First checks | Next controlled experiment |
|---|---|---|
| Out of memory before the first update | Other GPU processes, image resolution, wrong model family, precision | Batch 1; turn on checkpointing; avoid simultaneous validation pipelines |
| Out of memory only during samples | Number and size of generated samples, live training graph | Generate one image after training has exited |
| NaN or infinite loss | Data range, DC-AE output, component dtypes, learning rate | Reduce learning rate; isolate FP32 forward on a tiny batch; then inspect the exact failing component |
| Style appears but every image has the same object | Dataset/caption entanglement | Add independent objects and backgrounds; evaluate an earlier checkpoint |
| Trigger seems ignored | Wrong adapter loaded, weak/incorrect captions, insufficient variation | Verify adapter file and base revision; test paired fixed-seed prompts |
| Fine detail is always poor | Source resolution, crop damage, autoencoder reconstruction | Improve data and representation before extending training |
| Fast loss reduction but poorer generation | Overfitting, repeated data, excessive guidance | Compare earlier checkpoints and lower guidance separately |
The practical next-step ladder is:
- This Sana 1.6B BF16 LoRA configuration: batch 1, rank/alpha 8, 1,024-square images, checkpointing, frozen-component offload. The official script and its memory options support the approach; the book does not supply a measured VRAM number. Test loading, first backward, optimizer allocation, saving and separate inference before a long run D31 D36.
- Sana 600M: the official card lists a 590M transformer with a 512-pixel variant. It still needs its Gemma text encoder and DC-AE, so it is not a 590M complete application. It uses a different preferred precision/variant. Treat switching as a new checked experiment, not just changing the model name D37.
- Full transformer fine-tuning: optimizer states and gradients increase sharply when you stop freezing the backbone. Prove a need beyond the adapter baseline, reduce scale and measure before assuming it fits. A small complete model such as the audio project is a better first from-scratch exercise.
These are maintained practical choices rather than an exhaustive leaderboard. Older SD 1.x systems illustrate a U-Net and epsilon-prediction contrast, while Sana demonstrates a transformer and flow-matching target; their adapters, autoencoders, text encoders and samplers are not interchangeable. There is no detailed legacy training route in this package.
U Net versus DiT another architecture the same project questions
A U-Net reduces spatial resolution while building features, then upsamples while reusing earlier features through skip connections. This combines local detail with wider context. A diffusion transformer, or DiT, processes patches or latent tokens with transformer blocks. Attention lets positions exchange information, but large token counts increase computation and often memory. Neither name tells you whether the model predicts noise, clean data, a velocity parameterization or a flow field. Architecture and training objective are separate choices D06.
For practical adaptation, ask four questions before changing a command: What tensor enters the denoiser? What condition does it accept? What target was it pretrained to predict? Which parameters does this trainer actually update? The answers matter more than whether the model's name contains “diffusion.”