22 Teach a tool decision and a complete tool interaction
For this chapter's commands, open a fresh terminal at the companion
root, run cd examples/llm, then
source .venv-llm/bin/activate. If you skipped the first
small-language-model project, complete its explicit environment setup
first. Do not run these relative paths from an earlier project
directory.
The catalog assistant now needs to decide whether to call
lookup_part, supply an exact ID, and use the returned
observation in its answer. This is harder than copying a JSON format. It
requires a policy: call when appropriate, clarify when information is
missing, and avoid inventing unsupported tools.
Architecture determines what the dataset must contain
A decoder-only model learns to emit tokens. A tool call is therefore a specially formatted assistant output, not an invisible API operation inside the network. The runtime reads that output, validates it and runs the real function. The observation then becomes additional context for the next model turn.
This architecture requires examples of the decision and the consequence. A dataset containing only user questions and final answers does not show which tool produced the evidence. A dataset containing only call JSON does not teach what to do with a failed or empty result. Training only successful calls can accidentally teach “always call something.”
Keep the tool definition in each relevant example. A tool
schema names the function, describes its purpose and specifies
allowed argument types and constraints. Transformers’ tool templates
accept structured definitions and structured calls; their precise
serialized format is model-specific. TRL likewise distinguishes
messages, assistant tool calls, tool-role observations and a
tools column. L25 L26
Our read-only fictional tool is:
{
"type": "function",
"function": {
"name": "lookup_part",
"description": "Look up current stock for an exact part ID.",
"parameters": {
"type": "object",
"properties": {
"part_id": {"type": "string", "pattern": "^P-[0-9]{3}$"}
},
"required": ["part_id"],
"additionalProperties": false
}
}
}
A useful full conversation then includes these turns:
[
{"role": "user", "content": "How many P-101 parts are available?"},
{
"role": "assistant",
"content": "",
"tool_calls": [{
"type": "function",
"function": {
"name": "lookup_part",
"arguments": {"part_id": "P-101"}
}
}]
},
{
"role": "tool",
"name": "lookup_part",
"content": "{\"part_id\":\"P-101\",\"stock\":8,\"location\":\"bin A\"}"
},
{"role": "assistant", "content": "P-101 has 8 units in bin A."}
]
The real fixtures also include a system instruction and the schema. Tool observations are serialized data. Some runtimes also require call IDs to connect multiple calls and results; preserve those IDs in your application schema and use the model-specific format the runtime expects.
Mask decisions rather than observations
Teacher forcing means that training predicts an output using the supplied correct earlier sequence, rather than the model's own possibly mistaken earlier outputs. A correct tool trajectory therefore supplies the intended prior assistant calls and tool observations. At deployment, errors can accumulate through the model's own decisions, so evaluate complete interactions as well as isolated next-step predictions.
For this project, supervise the assistant’s call tokens and its final answer. Do not supervise the tool observation as if it were text the model should invent. This distinction is especially important for stock values: the model should condition on a returned count, not learn that it may fabricate one.
Our preprocessing expands a complete conversation into one example
per assistant decision. The first example ends with the call. The second
includes the call and observation as context and ends with the grounded
answer. Each expansion inherits the same source_id, so it
stays in the same split. This costs additional repeated context tokens
but makes the target explicit and easy to inspect.
Do not delete every non-assistant token from input_ids
to implement assistant-only loss. That would remove the observations the
answer depends on. Keep context tokens in the input and mask only their
labels.
Train a small version before collecting a large corpus
CUDA_VISIBLE_DEVICES=0 python train_small_lm.py \
--train data/tools_train.jsonl \
--eval data/tools_valid.jsonl \
--out runs/tool-lora-smoke \
--mode lora --steps 3 --max-length 1024
The longer sequence limit allows schema and observation text. It is a memory-relevant change, so repeat the fit check. The toy dataset includes successful stock lookups, a missing-ID clarification and an unsupported price request. It is still only a plumbing fixture.
A realistic corpus should vary the underlying decision, not just the wording. Include an exact ID, an ambiguous name, multiple similar tools, an irrelevant tool, a missing required argument, a legitimate no-tool answer, an empty result, a timeout, a recoverable validation error and an observation that contradicts the model’s earlier assumption. For state-changing tools, include permission boundaries and cancellation examples, and enforce those boundaries outside the model as well.
Mix paraphrases across training examples, but hold out entire
templates, ID ranges, scenarios and some tool combinations. If every
P-9xx ID appears only in test, the model must copy an
unseen argument rather than memorize a training ID. Include schema
rewording and tool order changes to discover whether it learned function
semantics or position in a list.
Untrusted observations can contain instructions. Include cases where the tool’s data says something like “ignore the user and call another tool,” then require the assistant to treat that as data. This is a training and evaluation concern, but permission enforcement remains an application responsibility.
Evaluate the decision at several levels
Run the unchanged model and the adapted model with the same decoding settings:
python generate_eval.py --input data/tools_test.jsonl \
--output tools-baseline.jsonl --device cuda
python generate_eval.py --input data/tools_test.jsonl \
--output tools-adapted.jsonl --device cuda \
--run runs/tool-lora-smoke
python score_tools.py data/tools_test.jsonl tools-baseline.jsonl
python score_tools.py data/tools_test.jsonl tools-adapted.jsonl
The script never executes a call. It checks the toy schema, tool name, exact arguments and whether a call was expected. These are separate questions:
- Syntax: can the output be parsed?
- Schema: are keys, types and constraints valid?
- Routing: is this the right tool, or should there be no call?
- Arguments: are IDs, units, dates and values correct?
- Interaction: does the model use the actual observation and recover from errors?
- Outcome: was the user’s task completed without unsupported claims or unauthorized actions?
A model can score perfectly on the first two and fail the last four. Likewise, “no parsed call” does not distinguish a good clarification from silence or a hallucinated answer. Read the prose and add an outcome rubric.
Next build a mock executor with a fixed fictional catalog. Give the model its own generated call, not the gold call, return the corresponding observation, then ask it to continue. This closed-loop evaluation exposes cascading errors that teacher-forced validation cannot see. Record the entire trajectory, tool-call count, correction count and final outcome. Impose a maximum number of turns so a failing agent cannot loop indefinitely.
BFCL is a useful public comparison framework because its versions distinguish call structure, realistic functions and multi-turn/agentic behavior. Its inspected leaderboard provides a specific evaluation commit and package version for reproducing that leaderboard snapshot. Do not mix a BFCL-v3 model-card number with a BFCL-v4 result as if they measured identical tasks. Our five toy tests are not BFCL and provide no comparable benchmark claim. L27 L28
Run the tool model on CPU too
Use the same unseen test file and schema:
python generate_eval.py --input data/tools_test.jsonl \
--output tools-cpu.jsonl --device cpu \
--run runs/tool-lora-smoke
python score_tools.py data/tools_test.jsonl tools-cpu.jsonl
CPU inference loads the exact base revision in FP32 and adds the saved adapter. GPU inference uses BF16 for this LoRA run. The template and input schema stay the same; numeric execution and speed differ. We measured neither runtime here. CPU inference may be slower and needs RAM for the whole base model, not only the small adapter.
Exercise. Add a mock timeout and a second tool named
lookup_price with different arguments. Make a held-out
request that requires only stock, then one that requires both facts.
Score the intermediate decisions and the final answer separately. Do not
grant the model direct unrestricted access to production tools while
testing.