Training Your
Own Models
Browse the book
Chapter 23 / 4010 min read

23 Study a small typed decision service

What this category does

A decision model answers a bounded question about supplied evidence. Its output is a label, a yes/no probability, a distribution over candidate actions, or a rating on a defined scale. It can route a ticket, shortlist a relevant file, flag a message for review, or choose which existing tool should handle a request. It does not have to generate prose to do those jobs.

The useful unit of application design is: evidence, question, allowed answers, estimated probabilities, and an explicit policy. The neural model estimates a judgment. Ordinary code validates the input, applies thresholds, checks permissions, and performs any authorized action. Returning a valid label solves a formatting problem; it does not prove that the judgment is true.

Start with a harmless project: suggest a support queue for fictional tickets. Do not start by giving a newly trained model authority to delete files, approve transfers, or make consequential decisions about people. Keep the first result as a recommendation shown beside the original ticket.

This chapter is an inference and architecture study of the published Laya interface, not a reproduced Laya training project. The following chapter supplies the complete local training, calibration, evaluation, and CPU/GPU inference workflow using an original Qwen-derived decision head. Treat the Laya snippet as a source-reviewed study example until its official environment and model download pass your local checks.

Why Laya is a useful first study model

The exact model here is convaiinnovations/laya. The root English checkpoint is approximately 421M parameters, using a fully fine-tuned ModernBERT-large backbone and an additional decision head. It fits comfortably below the book's one-billion-parameter example ceiling, although memory and runtime must still be measured for the chosen training settings. The multilingual and typed-decisions variants are distinct checkpoints with different intended uses and limits. J24

This gives a useful progression: first observe a trained classifier; then inspect its input and output shapes; then adapt it on a focused dataset; finally construct a similar interface using a causal language-model backbone. Small does not mean trivial: the data recipe, readout, calibration, and evaluation remain central.

Run the idea before studying all the internals

After using the project's official installation instructions in an isolated environment, a source-reviewed inference example is:

from laya import Router

router = Router(device="cpu")
result = router.predict(
    "The courier delivered my parcel to the wrong building.",
    {
        "queue": {
            "type": "choice",
            "instructions": "Which team should investigate this request?",
            "criteria": {
                "billing": "Payments, receipts, invoices, unexpected charges",
                "delivery": "Couriers, tracking, late or missing parcels",
                "access": "Passwords, sign-in, locked accounts",
                "review": "Insufficient information or a different problem",
            },
        }
    },
    model="english",
)
print(result["answers"]["queue"])

Record the installed Laya, Transformers, PyTorch, and tokenizer versions and pin the model revision in the supported loader configuration. The first call can download weights. This example has not been executed in the book's preparation environment. The repository's API and dependencies are moving, so inspect the pinned release's constructor and loader signature before execution. The inspected repository describes Python 3.10+ and separate CPU/GPU setup paths. J25

Now vary only one thing at a time: replace delivery evidence with a billing problem; remove the critical sentence; add irrelevant text; rename the option IDs while keeping descriptions; change the order of options. Record which changes should and should not change the answer. These probes tell you which architecture and training questions are worth studying.

What Laya reads and what the head computes

The shipped model code builds a sequence containing question type and instructions, an option marker before each option, and the state. Its encoder is bidirectional, so a marker before an option can receive information from later option text and later state text. After the encoder, a learned question-type vector is added, additional transformer layers refine the sequence, and a scorer maps each gathered option-marker vector to one logit. A separate head produces act/escalate logits. J26

For one illustrative batch with B = 2 questions, padded length T = 128, and D = 1024, the encoder emits [2, 128, 1024]. If each question has K = 3 options, gathering their marker positions produces [2, 3, 1024]. The option scorer maps that to [2, 3]. Softmax is applied independently across each row's three candidates. Padded candidate positions must receive no probability. The width-1024 value is grounded in ModernBERT-large's configuration; 128 and 3 are our teaching choices. J32

A key contrast with the later Qwen project is marker placement. In a causal decoder, a marker placed before the option cannot read the option's later text. We will instead read an option-end position and a final decision position. Copying the Laya layout into an unchanged causal decoder would remove information from the readout.

ModernBERT mixes local and global bidirectional attention. A small diagram that permits every token to read every other token is an intuition, not a complete implementation diagram of every ModernBERT layer. Read the backbone paper after the first inference probes make the role of attention concrete. J31

Training data is part of the architecture

Every training question needs a state, instructions, option text, and a target aligned with those options. Labels can be a single correct choice or an explicitly sourced target distribution. If a teacher model supplies probabilities, these are teacher beliefs, not automatically ground truth. A model can learn the teacher's blind spots.

For yes/no, keep a stable convention for which index means false and which means true. For ordered ratings, order is meaningful: shuffling levels without preserving their semantics invalidates an ordinal target. For ordinary categories, randomizing option order can discourage position shortcuts, but the target must be permuted with the options.

The shipped English configuration gives a maximum length of 512 and an option/question budget of 192. Those are distinct from the backbone's architectural maximum. Giving many lengthy labels the same finite option budget can truncate away the words that distinguish them. Inspect what the tokenizer actually keeps. J27

Treat changes to prompts, label descriptions, ordering, truncation, or tokenizer as changes to the input contract. Re-evaluate them as carefully as a changed weight file. A better input representation may improve quality more than increasing model size.

What the public training example really optimizes

The inspected single-process training script perturbs logits, scores the resulting probability distributions, forms a relative-reward policy-gradient term, and combines it with supervised cross-entropy. It updates encoder and head parameter groups at different learning rates and supports activation checkpointing. It also reserves a subset of encoded items for calibration. This is more precise than saying only that it uses RL. The script is written for MPS/CPU; do not paste its device command unchanged into an NVIDIA training guide. J28

A strictly proper scoring rule rewards reporting the true distribution at its ideal population optimum. That mathematical property is not a guarantee that a finite neural network, trained on limited or biased data, will be calibrated on a new domain. Clipping, finite optimization, sampling, and distribution shift also require empirical checks. Start a training comparison with cross-entropy, then add the more complex objective only if it helps a held-out metric.

The upstream Kaggle notebook is a useful public training walkthrough. The repository specifically warns about its calibration samples coming from training items. The separate current MPS script does reserve items, but questions from the same source can still leak across item-level splits. Split whole source records before expanding their questions. Keep model-selection, calibration, and final-test roles distinct. J29 J25

Read the limitations as carefully as the headline

The Laya card reports substantial zero-shot weakness on its typed-decisions test, gains from task-specific fine-tuning, overconfidence, label sensitivity, and an act_probability output with poor decision value. These findings make it a good study of why empirical validation matters. They do not justify treating every probability field as calibrated correctness. Its published comparisons with Jev also use third-party Jev results rather than a same-run controlled comparison. J24

For this project, ignore the act/escalate output until it passes a separate held-out evaluation. Fit any probability temperature on calibration data, choose the abstention policy on appropriate held-out data, and test both risk and coverage. Never promote a threshold such as 0.9 from a tutorial into a universal safety rule.

CPU and GPU are two deployments of the learned function

Inference does not need gradients, optimizer moments, or training activations. This makes deployment much lighter than full training. A 421M-parameter model has a raw fp32 weight-size estimate of roughly 1.68 GB, or roughly 0.84 GB at two bytes per parameter, before runtime buffers and duplicate copies. Those are arithmetic estimates. An actual loader can temporarily hold more memory.

CPU inference is useful for low-volume testing and small deployments, but throughput depends on CPU vector instructions, memory bandwidth, threading, and implementation. A GPU usually becomes more attractive for larger batches and repeated calls. Tiny batches can be dominated by transfers and launch overhead. Do not compare a warm GPU batch with a cold CPU model load.

For the book's custom Qwen pointer program, use CPU fp32 as the conservative portability path and CUDA bf16 autocast only when supported. For Laya and Kev, follow the pinned runtime's dtype/device support rather than assuming the custom program's choices are universal. An exported quantized model must preserve the custom head and formatting and pass probability-parity tests. Refit or verify calibration after a representation-changing export.

Useful application patterns

The user-provided Ten Levels of Jev repository is an application lab, not a training recipe or proof of Jev's internal architecture. Its original progression includes bounded classification, multi-question scoring, routing, confidence gates, file selection, and integration into a coding agent. It uses live hosted services and an explicit offline mock. The associated video is a useful walkthrough of that codebase. J33 J37

Here is a small, safe progression for your own trained model:

  1. Queue suggestion: evidence is one ticket; candidates are named queues plus review; output is a suggestion beside the evidence
  2. Multi-question triage: separately estimate queue, urgency, and whether required information is missing; deterministic code combines those answers into a worklist
  3. Selective automation: on a held-out set, choose a threshold that meets a stated error budget; send remaining items for review and measure the retained fraction
  4. Model routing: choose between a deterministic handler, a small text model, a larger reasoning model, and a person; evaluate final task success and total cost, including routing mistakes
  5. Retrieve, then judge: use an embedding model to shortlist documents, then a decision scorer to assess relevance; measure retrieval misses separately from judging mistakes
  6. Candidate verification: generate a small set of possible solutions, then rank them against the original requirements; compare chosen-solution success with random selection and the generator's first answer
  7. Agent observation: inspect tool results or change summaries and flag suspicious cases; preserve sandbox restrictions and independent authorization controls regardless of the model's confidence

These are application designs, not automatic evidence that any checkpoint is good at them. Build labels matching the intended decision, including failure cases. A candidate verifier cannot recover a correct candidate that the generator never supplied. A router can save model calls while lowering overall task accuracy if it sends hard examples to the wrong handler.

Keep the service boundary honest

The inspected Ten Levels client separates an explicit mock provider from live providers, validates requests, imposes a total timeout, limits retries, and records provider/model metadata. It distinguishes reported, estimated, and unknown cost. Those are practical integration ideas worth retaining when replacing hosted Jev with a local model. Passing mock tests validates application behavior, not model accuracy. J34

Its published confidence-gating example contains hardcoded thresholds, including one that permits destructive actions above a confidence bar. Treat that as a demonstration of branching logic, not a security design to copy. Model confidence must never replace permissions, least-privilege execution, deterministic path boundaries, confirmations required by the application, or human review for high-impact actions. J35

An API-compatible server is not a checkpoint-compatible model. Similar JSON fields across Laya, Kev, and CLM can carry differently defined confidence values, preprocessing, context limits, and performance. Build contract tests and a shared labeled evaluation before substituting one backend. Pin model IDs instead of moving latest aliases when results must be reproducible. J36

Reading and watching order

First run the harmless ticket exercise. Then watch the Ten Levels video with the repo open, inspecting how model outputs enter ordinary control flow. Next read Laya's implementation and the ModernBERT paper to explain the input and gathered-marker tensors. Finally study the Qwen pointer capstone and compare it to Kev. Karpathy's build-GPT video is a complementary code-first lesson on decoder construction; nanoGPT is useful historical source, but its README now marks it deprecated and points to nanochat. J37 J38 J39

The published Laya browser-agent adaptation is a later case study in how data and input formatting affect a working application. Its author-reported outcome is not our reproduction; use its data-generation and evaluation methodology as material to critique before attempting browser control. J30