28 Project build a tiny sound generator from random weights
Audio diffusion generates a waveform, an acoustic representation or a compressed audio latent by iterative refinement. Depending on its training and conditioning, it can create effects, ambience, music, or components of a speech system. A model trained on short sound effects should not be treated as a text-to-speech engine. If the task is transcription, classification, speaker verification or onset detection, choose a model trained for that discriminative or recognition task instead.
Useful application patterns include:
- Prompt plus duration → sound-effect candidates. Return WAV files with explicit sample rate, channels, duration, seed and model version. Audition candidates before placing them in a game or video; check event timing and tails, not just timbre.
- Class ID → a narrow family of variations. For example, generate alternative UI pings or percussion hits. This is the educational project below. A class ID is a finite choice; it does not understand arbitrary sentences.
- Existing audio plus a mask or context → repair/continuation. This requires a model explicitly trained for the input contract. Preserve timing, inspect transitions and make the provenance of generated regions visible. The tiny class-conditioned model below does not support this by itself.
- Acoustic features → waveform. A conditional diffusion vocoder can render mel features produced elsewhere. Feature extraction parameters and vocoder expectations must match exactly; this is not the same input contract as text-to-audio.
For an application, put generation in a job with a cancel path, store the requested duration and model settings, and return a playable file rather than an unlabeled tensor. Keep the original source recording and generated result separate. Apply a deliberate playback-level policy and never begin auditioning unverified output at high volume. Quality gates should include clipping, silence, audible artifacts, content adherence and consent/provenance. The first project chooses finite sound classes and very short mono waveforms so each part of the contract is inspectable.
What you will make. A small model that generates half-second thumps, pings and hisses when given a class label. The sounds are synthesized by our own code, not scraped recordings. This is a deliberately bounded audio task: it will not produce intelligible speech, human-quality music, arbitrary text-to-audio, or realistic room acoustics.
Why this follows the image adapter. You have already seen corrupt-data training and iterative generation. Here the whole neural network is trained, so there is no pretrained visual or audio knowledge hiding most of the work. The short waveform and simple classes make the experiment manageable and the failure modes understandable.
Audio project step 1 listen to the training problem
Return to the companion-script folder before this project. If you
opened a new terminal, set DIFFUSION_EXAMPLES to the actual
absolute folder path again. If the image virtual environment is active,
run deactivate first. Use a separate audio environment so
its libraries cannot silently replace the image project's
dependencies.
The commands below install the CUDA wheel for the intended GPU. For a
CPU-only smoke test, replace the PyTorch installation line with
python -m pip install torch==2.7.1 --index-url https://download.pytorch.org/whl/cpu,
and add --device cpu to training, sampling and evaluation
commands. CPU training is a mechanics check, not a practical substitute
for the planned GPU run.
cd "$DIFFUSION_EXAMPLES"
python3.11 -m venv .venv-audio
source .venv-audio/bin/activate
python -m pip install --upgrade pip
python -m pip install torch==2.7.1 --index-url https://download.pytorch.org/whl/cu126
python -m pip install numpy==1.26.4 scipy==1.15.3
python -m pip check
python make_tiny_sfx.py --out data/tiny_sfx --per-class 800
You now have 2,400 mono WAV files: 1,920 training, 240 validation and 240 test. Open several files at low playback volume. You should hear clearly different families, with variations in frequency, onset, decay and amplitude. If you cannot identify the intended classes reliably, fix the data before asking a model to learn them.
Waveform sample rate channels and duration
A waveform is a sequence of amplitude samples over
time. A sample rate of 16,000 Hz means 16,000 amplitude
measurements per second per channel; it is not the number of training
examples. Mono has one channel; stereo has two. For
this project the tensor shape for one example is [1, 8192]:
one channel, 8,192 values. Its duration is
8192 / 16000 = 0.512 seconds.
The file stores signed 16-bit PCM. Loading converts those integers to
floating-point numbers approximately in [-1, 1]. “16-bit
audio” describes file quantization; “FP16 training” describes arithmetic
precision inside the model. They are different decisions.
Changing only a WAV header from 44.1 kHz to 16 kHz changes playback speed and pitch. Proper resampling creates a new sequence with filtering. Frequencies above half the destination sample rate need suitable attenuation to avoid aliasing. Our generated data are created directly at 16 kHz, so the first project does not need resampling at all.
Longer audio is a real size increase. A 10-second stereo 44.1 kHz waveform contains 882,000 values, about 108 times as many as this mono half-second example. It is not a harmless setting change. In an attention model the token-cost increase can be even less forgiving.
One manifest record is:
path,label,split,group,source,rights,sha256
ping/ping_0000.wav,ping,train,ping_0000,synthetic_book_generator_seed_123,generated_locally_no_recorded_people,<actual checksum>
The actual generator writes the checksum. Preserve source and rights fields when replacing synthetic sounds. A real recording session, speaker, musician or original long file may define a group. All overlapping windows from the same original belong to the same split. Otherwise the validation set may be a slightly shifted version of training audio.
Audio project step 2 the smallest useful full training loop
The complete implementation is tiny_audio_ddpm.py. First
test one update, saving and loading:
python tiny_audio_ddpm.py train --data data/tiny_sfx \
--out runs/audio-smoke --steps 1 --batch 1 \
--log-every 1 --save-every 1
python tiny_audio_ddpm.py sample --checkpoint runs/audio-smoke/step-1.pt \
--class-name ping --n 1 --out samples/audio-smoke
The resulting sample should not be expected to sound good. Its
purpose is to prove that model input and output shapes match, the
checkpoint is readable and the sampler produces a valid WAV file. A
CPU-only mechanics check can use --device cpu; sampling
still invokes the network hundreds of times and may be slow.
Then run a measured pilot:
python tiny_audio_ddpm.py train --data data/tiny_sfx \
--out runs/audio-pilot --steps 200 --batch 4 \
--log-every 50 --save-every 200
Inspect metrics.jsonl, which records training MSE,
fixed-noise validation MSE, elapsed seconds, and CUDA peak
allocated/reserved GiB when using CUDA. The code prints the actual
parameter count. It uses float32 rather than hiding numerical details
behind a mixed-precision wrapper; the model is deliberately small. On a
24 GB card this is a conservative candidate workload, not a GPU-memory
measurement made by the author.
If it runs cleanly, continue the same run to a 10,000-update experiment:
python tiny_audio_ddpm.py train --data data/tiny_sfx \
--out runs/audio-pilot --resume runs/audio-pilot/step-200.pt \
--steps 10000 --batch 4 --log-every 100 --save-every 1000
The checkpoint includes network, exponential-moving-average network,
optimizer, diffusion-step count, random-generator state and a material
run configuration. Every WAV is checked against its manifest SHA-256
before use. Resume requires the original output directory, its latest
saved checkpoint, unchanged data, source code, batch, accumulation,
learning rate, seed, training device, numerical precision, schedule and
library versions. Increasing the total step target or changing
log/checkpoint frequency is allowed, as in this pilot-to-long example.
To change a material training setting, start a deliberately new
experiment. CPU inference remains possible for a CUDA-trained
checkpoint; this strict same-device rule applies to continuation of
optimization. The initial run.json is retained rather than
overwritten. Each continuation appends an invocation record with its
actual optimizer learning rate; metrics include an invocation ID, so
replayed steps after a crash are distinguishable. A fresh run refuses a
nonempty output directory. The supplied file format is intended for
checkpoints you created or trust; it uses
torch.load(..., weights_only=True) and should not be
changed to unrestricted loading to silence errors from an unknown
file.
Read the audio training loop in ordinary language
This time we use DDPM noise prediction, rather than Sana's flow
target. For clean waveform x0, Gaussian noise
epsilon and a timestep with retained signal power
alpha_bar[t], the corruption is:
xt = sqrt(alpha_bar[t]) * x0 + sqrt(1 - alpha_bar[t]) * epsilon
training target = epsilon
The square-root factors control signal and noise power. The network
gets xt, time and class, then estimates the actual added
noise. This differs from predicting epsilon - z0 along
Sana's straight path. Both produce a generator, but their loss targets
and reverse updates must not be mixed D01.
For every optimizer update, the script performs these operations:
- Select a small batch of clean waveforms and class IDs.
- Draw a timestep independently for each waveform.
- Draw fresh Gaussian noise of exactly the same shape as the waveform.
- Mix clean audio with noise according to the selected noise level.
- Drop the class label for roughly 10% of examples, substituting a learned null label.
- Ask the network to predict the added noise.
- Compare prediction and noise with mean squared error, backpropagate, clip gradients, and update the weights.
- Update a slowly moving copy of the weights for evaluation and sampling.
The class label is conditioning; the target is still noise. The training pairs are not “hiss goes in, ping comes out.” At high noise levels all three classes look noise-like. The condition helps identify which clean distribution the reverse process should favor.
The model is a small 1D U-Net: convolutions operate along time instead of across image width and height. It reduces the temporal length by factors of four three times, processes the compressed features, then upsamples while using skip connections. Time and class embeddings modify intermediate features. This is an original educational implementation, not the released DiffWave architecture. DiffWave is a primary example of applying diffusion directly to waveforms D18.
Why these audio hyperparameters are small
- 8,192 samples: enough for a short transient but not for phrases, melodies or long room reverberation.
- Batch 4: a conservative initial batch. Reduce to 1
if you need to diagnose memory; use
--accum 4if you want roughly the same effective batch afterward. - Learning rate 0.0002: a starting value for this newly initialized small network. It is not a recommended rate for an arbitrary pretrained audio transformer.
- 256 diffusion levels: the discretization used by both training and the simple ancestral sampler. It does not mean 256 optimizer steps.
- Cosine noise schedule: brings the terminal signal close to pure noise while distributing intermediate signal levels. The cumulative signal-power curve is converted into per-step noise amounts. This follows the schedule idea in Improved DDPM rather than silently reusing a thousand-step linear schedule at 256 steps D17.
- 0.1 condition dropout: teaches a null-conditioned case, enabling ordinary classifier-free guidance later.
- EMA factor 0.999: the evaluation copy retains most of its previous value and incorporates 0.1% of the current weights each update. Early in a short run it lags considerably; do not judge a one-step EMA sample as a trained model.
- 10,000 updates: a bounded experiment. At batch 4, it is 40,000 waveform presentations, about 20.8 average passes over the 1,920 training examples. Training uses random sampling with replacement, so these are average exposures, not exact complete epochs.
The code does not include every production optimization. This is useful: you can inspect the loss, forward noising and reverse update without disentangling several frameworks. If the experiment is too slow, measure where time goes before adding mixed precision, compilation, larger batches or a different sampler.
Audio project step 3 generate and compare every class
for CLASS in thump ping hiss; do
python tiny_audio_ddpm.py sample \
--checkpoint runs/audio-pilot/step-10000.pt \
--class-name "$CLASS" --n 16 --seed 17 --cfg 1 \
--out samples/audio-10000
done
python audit_audio.py --data data/tiny_sfx --samples samples/audio-10000
python tiny_audio_ddpm.py evaluate --data data/tiny_sfx \
--checkpoint runs/audio-pilot/step-10000.pt --split val
The sampler starts with a new noise waveform, evaluates the denoiser at decreasing noise levels and applies the matching DDPM posterior update. Its final clean estimate is bounded to the waveform range. The file writer preserves generated relative amplitude instead of peak-normalizing every clip. That makes loudness and clipping defects visible in the diagnostics.
At first, use --cfg 1, which gives the conditional
prediction without extrapolation under our formula. Then make a separate
comparison with --cfg 2. Increasing guidance must not be
the only way to make classes recognizable; it can amplify harshness and
reduce variation. The two prediction calls used by guidance add
inference work. They do not retrain the model.
Listen quietly first. A bad generator can output unexpectedly loud noise, tones or clicks. Never judge sound quality only by looking at an attractive spectrogram.
Audio inference on GPU and CPU
Training and inference use the same class order, waveform length and sample rate. The checkpoint records these values and the cosine schedule identity; the loader rejects mismatches. Inference restores the EMA weights, creates the matching schedule and changes only the device. Both CPU and CUDA paths in this small implementation use float32. There is no tokenizer, text encoder, latent codec or vocoder to substitute: conditioning is the stored class vocabulary, and the generated tensor already is a waveform.
To generate one sample on CPU:
python tiny_audio_ddpm.py sample \
--checkpoint runs/audio-pilot/step-10000.pt \
--class-name ping --n 1 --seed 17 --cfg 1 --device cpu \
--out samples/audio-cpu
CPU generation is implemented but still requires 256 network evaluations per unguided sample; guidance above 1 uses two predictions per step. It is appropriate for a functionality check or occasional output, not a claimed real-time engine. GPU and CPU outputs need not be numerically identical. The writer clips to the valid range and quantizes to mono PCM16 at 16 kHz; it does not resample or peak-normalize each sample. Thus the saved half-second duration and relative amplitude retain the training contract. For evaluation on either device, the real-input loader uses the same PCM-to-float conversion as training.
The source-code hash and data-manifest hash are stored in training checkpoints. Retain the code and manifest that produced a checkpoint. Increasing duration, changing class order or swapping the noise schedule without retraining is not a portable inference optimization.
An audio evaluation sheet that answers the task
For each checkpoint, generate the same class/seed combinations. Rename or shuffle files before listening so that knowing the checkpoint does not bias your score. Evaluate:
- Class correctness: can a listener distinguish thump, ping and hiss without seeing the filename?
- Signal integrity: unexpected DC offset, near-full-scale samples, clicks at boundaries, persistent noise, abrupt truncation or unwanted silence.
- Diversity: varied valid frequencies, decays and onsets within each class. Listen to several outputs in sequence.
- Distribution match: compare ranges of duration, amplitude, spectral centroid and envelope with held-out real/synthetic examples. A model that always emits the loudest ping has not captured the distribution.
- Memorization: compare suspicious samples with nearest training sounds. The included audit uses log-magnitude short-time Fourier features to screen for spectral neighbors. It is a review aid, not proof that an output is or is not a copy.
MSE on newly noised held-out waveforms is a useful learning diagnostic, especially with fixed noise and timesteps for checkpoint comparisons. It is not a listening score. Waveform MSE between two independently generated sounds is usually not a meaningful quality metric: a tiny phase shift can produce a large pointwise error while sounding similar.
A spectrogram shows energy at frequencies over time, computed from short overlapping waveform windows. A mel spectrogram pools those frequencies on a perceptually motivated scale. Both are representations, not automatically playable audio. A spectrogram-generating model needs a reconstruction process or vocoder, and its settings must match that decoder. A vocoder turns acoustic features into a waveform; it is another model or algorithm to train, acquire and evaluate. Our project uses waveforms directly so there is no hidden vocoder.
For larger text-to-audio work, embedding-based audio quality/distribution measures and audio-text similarity may be informative when computed with a documented protocol and enough examples. They can be fooled by background cues or domain mismatch. Do not use a single audio-text similarity score as a substitute for checking clicks, intelligibility, rhythmic continuity or whether a prompted event actually occurs.
Only after choosing the checkpoint using validation should you run:
python tiny_audio_ddpm.py evaluate --data data/tiny_sfx \
--checkpoint runs/audio-pilot/step-10000.pt --split test
If an earlier checkpoint is better, change the checkpoint path. The objective is a defensible selection, not necessarily using the largest step number.
Audio project step 4 replace synthesized clips carefully
A safe next experiment is a narrow collection of your own non-vocal sounds: keyboard clicks, hand percussion, doors, small tools, or foley props. A directory of random songs is a much harder task with different rights and temporal structure.
Start by preparing files offline to the exact format the educational
loader accepts: mono PCM16 WAV, 16 kHz, exactly 8,192 samples. The
loader intentionally rejects other formats so that a sample-rate mistake
cannot silently train the wrong task. Create a new manifest and matching
class set; editing CLASSES changes the model's label
vocabulary and makes old checkpoints incompatible.
For real recordings:
- Keep the original audio unchanged, with provenance and permission records.
- Decode to floating point and inspect channel layout. Do not blindly average stereo: oppositely phased channels can cancel.
- Resample with an actual resampler. Document the original and final sample rates.
- Choose event-centered windows for transients. Random cropping may turn a useful sound into silence.
- Pad only when necessary, track true length, and decide whether padding should contribute to the loss. The tiny fixed-length example assumes every position is part of the example; a serious variable-length trainer needs masks or other explicit length handling.
- Remove corrupt and heavily clipped examples. Do not normalize silent clips by dividing by a near-zero peak.
- Use sensible common gain or a documented loudness policy. Per-clip peak normalization can erase meaningful loudness variation and raise background noise.
- Split by original recording/session before extracting windows, and listen to the final processed files.
A half-second mono 16 kHz model has a limited frequency range and context. If your target includes cymbal shimmer, stereo space or several seconds of decay, acknowledge that this representation is insufficient. Choose the correct duration, bandwidth and channel layout before a long run, then remeasure memory.
For text-conditioned datasets, a sidecar or manifest caption should describe the actual audible event, texture, environment and sequence, not tags copied from an unrelated image. “Three close dry wooden taps, with a short pause before the last” conveys more useful structure than “excellent quality sound.” Captions cannot make timing controllable if the architecture never receives timing information or the examples contradict it.
Voice music and dataset provenance
Recorded voices may identify people even when filenames do not. Obtain explicit informed permission for the intended voice-model training and generation, including scope of use and retention, and be able to separate or delete that person's source records when required. Owning a recording device or finding a public video does not establish permission to clone the speaker. Do not impersonate someone or imply that they said generated words.
For music and effects, the recording, composition, performance and dataset license can involve different rights. Record the source URL or original filename, creator, exact license/version, attribution requirements, acquisition date, transformations and intended use. A repository's software license does not automatically license its training audio. Model licenses are a separate layer. Inspect the actual current terms rather than treating “Creative Commons,” “open,” or “research” as interchangeable legal categories.
Synthetic examples in this project avoid recorded voices and borrowed artwork. That makes the first experiment easier to audit; it does not establish blanket rights for future datasets or guarantee that every pretrained model's outputs are risk-free.
After the tiny project pretrained audio adaptation
A realistic text-to-audio system commonly uses a pretrained waveform autoencoder, a text encoder and a latent denoiser. This parallels the image pipeline, but the compression rate, frequency fidelity, stereo handling and long-range temporal structure are different. Before adapting the denoiser, test autoencoder reconstruction on your target sounds: metallic transients and unusual textures can reveal representation failures.
Stable Audio Open 1.0 is an established official
option: its card describes text-conditioned stereo generation at 44.1
kHz up to 47 seconds, using an autoencoder, a T5-based text
representation and a latent DiT. The official
stable-audio-tools repository contains training machinery,
but a model-card inference snippet is not by itself a verified
fine-tuning configuration D19 D21.
Stable Audio Open Small is a different model: its card specifies up to 11 seconds, stereo 44.1 kHz, and a short distilled inference recipe. Do not assume that ordinary noise-prediction fine-tuning preserves the behavior of a distilled model or that any LoRA trainer supports it. Check the exact model, objective, sampler and trainer together D20.
Stable Audio 3 now has an explicit official
adaptation path. Its repository distinguishes base checkpoints for
training from post-trained inference checkpoints, and the inspected
train_lora.py accepts only base-model names. It supports
raw audio with same-stem .txt captions or pre-encoded
latents, and exposes duration, batch, rank, learning rate and local CSV
logging. This is a genuine current route to investigate, rather than an
invented Stable Audio Open LoRA flag D22 D24.
A cautious upgrade procedure is:
- Use a separate checkout/environment. Read the official model license and download-access terms. Stable Audio 3 Small SFX also names Gemma terms for its text-conditioning component; a single umbrella “open weights” label is insufficient D26.
- Record the repository commit, model revision and resolved
uv.lock. The inspected project pins PyTorch 2.7.1; its own dependency environment is substantially different from our small script D25. - Inspect
python scripts/train_lora.py --helpat that commit and select one of its actual base-model names. Do not substitute an inference model just because its filename is similar. - Start with a few seconds of audio, batch 1 and a small adapter, and keep demonstrations bounded. In the inspected script, the demonstration callback uses the model configuration's sample size, which can differ from the training crop duration. A short training crop alone is not proof that demo generation will be small.
- Run a very short optimization-and-reload test before budgeting a full run. Measure preprocessing, first backward pass, optimizer state allocation, checkpointing and demonstration peaks separately.
- Compare generated audio against the unchanged base under identical prompts, durations, seeds and sampling settings. Test content outside your adaptation domain to detect degradation.
Do not promote uncertain VRAM figures into a guarantee. Stable Audio 3's README labels its table as inference measurements on an H200. Its LoRA guide contains lower approximate memory entries but also a “medium on about 16 GB” example without a uniform duration/batch measurement protocol. These do not establish the peak VRAM of your 24 GB training run. The guide confirms support, not a benchmark reproduced in this book D22 D23. The tiny waveform project above is supplied in full precisely so your first audio training experiment does not depend on an unverified pretrained-model stack.
Revisit the theory noise prediction score prediction and flow matching
Now that you have trained both an adapter and a full small network, the terminology should connect to code rather than remain a list of names.
Noise prediction asks for the particular Gaussian
noise used to corrupt an example. Clean-data prediction
asks for the original clean tensor. A velocity
parameterization, often named v_prediction,
predicts a schedule-dependent combination such as
v = alpha * epsilon - sigma * x0, with
alpha² + sigma² = 1 for this convention. These targets can
be converted under the corresponding schedule, but their raw outputs are
not interchangeable. The training target, model output and sampler
interpretation must agree.
A score is the gradient of the log density of noisy
data with respect to the noisy tensor. It points toward higher density
locally; it is not a human quality score. For Gaussian corruption with
noise standard deviation sigma, the optimal noise predictor
and score have a known scaling relationship,
score ≈ -predicted_noise / sigma. Score-based formulations
describe reverse stochastic or deterministic dynamics using this
quantity D02.
Flow matching trains a time-dependent vector field
along a chosen probability path. A simple illustrative path from noise
z at time 0 to data x at time 1 is
x_t = (1-t)z + t x, with conditional target velocity
x-z. The network predicts a movement vector and an ODE
solver follows it. Real systems may use different paths, weightings,
time parameterizations and coupling schemes. Flow matching includes a
wider family of paths than this one example; it is related to diffusion
but is not merely renaming an epsilon-prediction loss D03.
Be especially careful with the word “velocity”: diffusion
v_prediction and a rectified-flow vector field can use
different definitions. Copying a loss from one into the other while
retaining the original sampler is a silent conceptual bug.
There are at least three different schedules in a training project:
- Learning-rate schedule: how optimizer step sizes change during training.
- Noise/path schedule and timestep sampling: how corrupted examples and targets are constructed and which regions of the path receive training weight.
- Sampling schedule: which noise/time levels and solver updates are used to generate an example after training.
Fewer sampling steps can speed inference without reducing the number of learned parameters. It can also reduce quality or require distillation. Changing the training noise schedule for a pretrained checkpoint is not a free speed optimization; it changes the problem the network sees. Start with the base model's objective and schedule, and change one component only when you can explain and test the consequences.
End of project checklist
You are ready to move on when you can answer these questions without memorizing a model list:
- What is the shape and numeric range of a clean example?
- What condition is given to the network, and what exact target does its loss compare against?
- Which components are frozen and which are optimized?
- What does one optimizer step mean, and how many examples does it represent?
- Which noise schedule and sampling rule interpret the network's outputs?
- What can the validation experiment establish, and what remains untested?
- What are the source rights and consent for the data?
- What is the measured peak memory of the actual configuration, including evaluation?
- Can you load a saved artifact in a fresh process and reproduce the evaluation setup?
- Does the model do something useful beyond reproducing training examples?
If you can answer those, the next model family is a controlled engineering change rather than a new collection of unexplained flags.
Read a public training process after building your own
Open the original DDPM repository beside the paper and find the functions that construct noisy data, define the prediction target and update a sample. Its TensorFlow 1.15/TPU environment is historical; use it to understand the published process rather than replacing this book's environment with its old requirements D28. Then compare the DiffWave repository's waveform and mel-conditioning paths with our finite class labels. Trace where its conditioning enters, what its preprocessing saves and how inference obtains a waveform D29. The exercise is to identify the data-model-loss-sampler contract in real code, not to assume that a paper's released configuration will fit your card.