21 Make a writing style reproducible without changing the facts
For this chapter's commands, open a fresh terminal at the companion
root, run cd examples/llm, then
source .venv-llm/bin/activate. If you skipped the first
small-language-model project, complete its explicit environment setup
first. Do not run these relative paths from an earlier project
directory.
Return to the concise-update project with a better dataset. You now know the mechanics; the main problem is defining what counts as the desired behavior.
Separate voice from subject matter
A decoder model can imitate correlations in its training examples. If every example of your desired voice discusses software releases, it may learn release vocabulary rather than a style that transfers. If every “concise” answer omits uncertainty, it may learn to sound confident instead of brief.
Use examples that contain both source content and the desired rewrite. Preserve the distinction between facts, requests, opinions and promises. Label uncertain facts as uncertain in both sides. Include examples where being concise still requires two paragraphs or a caution. A fixed length target can reward harmful omissions.
For personal style, use writing you own or are authorized to use. Remove secrets and unnecessary third-party details. Keep recipient/channel context when it affects wording, but do not teach the model to invent that relationship. “A short note to a colleague” and “a formal external update” may be separate style settings in one dataset.
The existing script fits this architecture-data relationship: source facts and the requested style are input context; the approved rewrite is the supervised continuation. A collection of standalone essays without the input conditions is a raw-text language-modeling dataset, not the same conditional rewriting task.
Build paired examples and a blind rubric
Each source item can have a rough version, an approved rewrite and a short explanation for a human annotator. The explanation need not be part of the model’s training target. Keep it as provenance: why was this a good answer?
Useful rubric dimensions are:
- Every necessary fact preserved
- No new claim, deadline, opinion or commitment
- Appropriate certainty
- Requested tone and channel conventions
- Clear structure and readable length
- No copied private or unrelated material
Evaluate on new topics and different source lengths. Randomize whether the baseline appears as answer A or B, hide model identities and allow ties. Judge content fidelity before stylistic preference. If a model judge helps scale evaluation, check a human-reviewed subset and inspect its reasons; it may prefer verbosity, familiar phrasing or its own writing habits.
Simple automatic checks can verify numbers, required names, forbidden boilerplate and maximum length. They are guardrails rather than a complete style score. A shorter sentence is not necessarily better. A high lexical match to the target is not necessarily better than a faithful paraphrase.
Start with prompting plus examples, then LoRA, and compare full fine-tuning only if the adapter experiment suggests a capacity or behavior limit. Keep the test set untouched. A stylistic win accompanied by more invented facts is a regression for this task.
Use the saved artifact on a genuinely new input
python generate_eval.py --run runs/style-lora-smoke \
--device cuda --max-new-tokens 96 \
--prompt "Write a concise update using only these facts: 19 checks passed; one failed; no retry is scheduled." \
--output fresh-style-gpu.jsonl
Repeat with --device cpu and a different output
filename. The script loads the manifest’s exact base revision, attaches
final/, loads the saved tokenizer, applies its chat
template, moves tensors to the selected device and generates under
inference_mode(). It slices off the prompt token IDs and
decodes only the new tokens. It does not accidentally report the
original prompt as generated output.
For a full-update run, use --run runs/style-full-smoke:
the script loads the complete final/ model directly. For
QLoRA on CUDA it reloads NF4; on CPU it deliberately uses the original
FP32 base plus adapter. The latter is a portable numerical variant, not
an assertion that the CUDA quantization backend works on your CPU.
Re-evaluate it. An architecture-supported CPU quantized export is
another deployment project, with its own conversion and quality
checks.
Exercise. Create three held-out inputs containing an exact number, an unresolved decision and a statement the writer has explicitly declined to promise. Have another person score content fidelity before style. Record examples where a fluent answer is still wrong.