Before training
Starting modelI’m not sure which team can help with this. A person should review your message.
Still needs a human to sort it out.
Fine-tuning, from the inside
The same conversation.
A model that’s learned what to look for.
I’m not sure which team can help with this. A person should review your message.
Still needs a human to sort it out.
That’s a delivery issue. The delivery team is the right place to help.
Finds the right team: Delivery
Same message. A more useful response.
Illustrated replies based on a tiny classifier’s real predictions.
Juniper is a fictional shop; this is not a live support chat or an LLM.
Imagine a chef who already knows how to cook. You want them to work in your kitchen.
You don’t teach them what a pan is. You give them examples of your menu, your portions, and how you plate each dish. Then they practice. That is the useful part of the original article’s chef metaphor: adapt an existing skill, rather than build everything again.
For a language model, the "practice” is mathematical. Training adjusts numbers called so the model is more likely to produce the kinds of answers in your examples. It can improve a particular behavior, but it can also learn mistakes or lose useful behavior. It does not become an infallible expert.[1][3]
Pre-training taught it broad language patterns. Your fine-tune begins from those learned weights, not random numbers.
In supervised fine-tuning (SFT), you supply a prompt and the desired answer. The model assigns probabilities to the next token, a small piece of text. The training program compares those probabilities with the target tokens and calculates a mismatch called loss. Gradients tell the optimizer how to change the trainable parameters.[2]
For next-token training, the model is normally given the preceding correct tokens while predicting the next one. It is not necessarily writing an entire fresh reply and receiving a human score after every example. In our practical recipe, the loss is applied to the target completion, not the user’s question.
predict → measure mismatch → adjust → repeatLower training loss means a better fit to those targets. It does not prove the targets are good.A simplified token illustration, not the real Qwen tokenization. Suppose the target next token is "billing.” Move the slider to change the probability assigned to it.
This is the actual cross-entropy calculation for one target. The slider changes a hypothetical probability; it does not run an LLM.
First ask: does the model need better instructions, better information, or more practice?
Tell it what to do in this request.
"Return only a category.”Find relevant information and include it in the request.
"Here’s our current policy.”Train it on examples of the behavior you want.
"Practice how we categorize.”These tools can work together. A tuned model can still receive instructions and retrieved evidence. For changing facts, retrieve a current source; for a live order status, call an authorized order service. Training is not a substitute for that lookup.[4]
Keep a baseline: the existing model, your best instructions, and examples in the prompt. Fine-tuning earns its place only when the comparison shows a useful improvement for your task.[3]
A training example says: "When the input looks like this, here is the output we want.”
Meet Juniper, our fictional online shop. We’re teaching a model to route a message to delivery, billing, account, or a human. Not to issue a refund. Not to access a customer record. Just one small, testable job.
Choose one team.
Reply with JSON only.
"You charged me
twice.”
{"queue":
"billing"}
A chat dataset stores these messages with role labels. The training software uses the chosen model’s to turn them into the token sequence it expects. The exact formatting matters.[5]
Use fictional or appropriately cleared data. Don’t paste real secrets or customer details.
Checks this exercise’s text-message format, allowed labels, exact duplicate inputs, conflicts, and a few sensitive-data patterns. It does not verify truth, rights, consent, or every provider’s upload rules. Review the actual examples.
The source article uses these three categories. General data covers broad topics; domain-specific data focuses on a field or task; mixed data combines them. Keep those terms, but don’t choose by the label alone.
For Juniper, we propose representative shop messages with reviewed routing labels. Include spelling mistakes, short messages, ambiguous requests, and all four teams. A pile of unrelated articles does not directly demonstrate the routing behavior. A mixed set may help preserve other capabilities, but that is something to measure, not a guaranteed benefit.
Raw domain text can be used for continued language-model training; instruction-and-answer examples teach a different, more explicit behavior. This guide’s runnable path focuses on the latter.[2]
This little model really changes its weights, right here in the page.
It starts with practice on 16 generic examples, then you adapt it with 16 shop messages. Try the training example "My mug arrived shattered.” before and after. Then change the label quality and watch the separate validation results on phrases it has not practiced.
This is a tiny word-based classifier with four outputs, not a language model. The plots show its actual loss and accuracy on this small, fixed dataset. Its speed, learning rates, and results do not predict an LLM fine-tune. The dataset editor above is a separate exercise; this lab uses its own fixed examples.
Hover over the chart or move the slider to inspect a round. Solid line: practice. Dashed line: validation.
The broken-mug preset is a training example, not an independent test. These probabilities describe this toy classifier. They are not a calibrated guarantee. It does not understand word order, negation, or every possible topic.
Each known word has one weight for each team. Words in a message become a normalized presence vector; a linear layer and softmax turn that vector into four probabilities. The bias is trainable too. Four examples are used per update.
Cells show current weight minus starting weight. This display is a real weight change, not a claim that LLM concepts live in individual labeled cells.
loss = −log(probability of the target team)
new weight = old weight − learning rate × gradientThe program uses mini-batch gradient descent. There are no lookup tables of "after” results.Validation examples do not update weights. They only measure the run. Download the trace to inspect every round.
How many times you practice the set.
How far each update moves.
How many examples contribute to an update.
More practice is not always better. A model can memorize quirks in the training set and do worse on new inputs. That is . Watch the held-out results, not just a reassuring downward training curve.[6]
This is the teaching target, not a new instruction.
My mug arrived shattered. → Delivery
Press Train one step. The model will compare its prediction with your label, then adjust its weights.
Real updates in a tiny word classifier. Each dot is an actual weight; its 3D position is just a visual layout. These are training-example probabilities, not LLM results or proof of accuracy.
Open the full training desk ↗Practice questions and final-exam questions must not be the same.
Used to change the weights.
Used to choose settings and checkpoints.
Kept aside for the final comparison.
Reserve the splits before tuning. Keep related records, such as messages from the same support conversation, together. A different sentence from the same thread can still leak the answer. Repeatedly choosing changes from the final-test results turns that test into another validation set.[7]
A fixed illustration with four fictional conversations, not a recommended train/validation percentage. In production, also check near-duplicates and consider time-based splits for future-facing tasks.
For this routing project, our proposed review separates valid JSON, correct team, and appropriate human escalation. Review results per team so a common category doesn’t hide failures in a rare one. Compare the existing model and the tuned model with the same prompt, decoding settings, and held-out cases.
For a writing assistant, we would add a human rubric for clarity, tone, factual support, and prohibited claims. Loss alone does not answer those questions. Test older capabilities too: specialization can introduce regressions. The supplied article calls this risk catastrophic forgetting.
You don’t always need to update the whole model.
Full fine-tuning updates the model’s parameters. LoRA freezes the original weights and learns smaller add-on matrices. Think of a small, trained adjustment layer, not a separate database, and not a complete replacement model.[8]
Tap the original matrix or small update to see which part can learn. This diagram illustrates the method; it does not run LoRA.
For one d × d matrix, full tuning updates d² parameters. A rank-r LoRA update adds d × r + r × d parameters. A full transformer has many matrices and other parameters, so this percentage is not the fraction for an entire model. Rank does not change the dimensions of the original matrix.[8]
QLoRA stores the frozen base at lower precision and trains adapters. "4-bit” describes the quantized base representation, not every intermediate calculation or all training memory. Activations, optimizer state, adapters, and quantization metadata still use space.[9][10]
SFT: "Here is the answer I want.” Train on demonstrations. DPO: "For this prompt, this answer is preferred over that one.” Train with chosen/rejected pairs. DPO is a different training objective, not just another name for LoRA.[11]
LoRA describes which parameters are adapted; SFT or DPO describes what learning signal is used. Our starter combines supervised learning with LoRA.
A small, complete project is better than a mysterious "Train” button.
The starter below adapts Qwen/Qwen2.5-0.5B-Instruct with LoRA using Hugging Face Transformers, TRL, and PEFT. It is a compact instructional example, not a claim that this is the best model for your business. The publisher’s model card documents the architecture, license, and usage.[12]
Define the job. One message in; one routing JSON object out.
Prepare & separate. Review the data; reserve validation and final-test files.
Record the baseline. Generate answers with the starting model and the same prompt.
Run SFT + LoRA. Train, save checkpoints, and select using validation loss.
Compare and release carefully. Measure held-out behavior; keep the old version available.
This is a hypothetical per-token billing model. Actual providers may count tokens differently, charge minimums, or charge for GPU time instead. The open-weight starter consumes compute rather than calling a per-token training API. Include data preparation, validation, storage, evaluation, retries, and inference in a real budget.
Effective batch = per-device batch × accumulation × devices. The displayed optimizer steps assume fixed-size examples, no packing, and a final partial batch counted each epoch. Framework details and distributed sampling can change the exact step count. More accumulation is a way to combine smaller micro-batches; it does not make compute free.
Count tokens using the chosen model’s tokenizer. The sample value of 300 is an assumption, not a conversion rule for words.
Save and unzip the starter files. Use Python 3.11+ and a compatible PyTorch installation. A CUDA GPU is the intended training path; CPU mode is available for a slow smoke test. The first run downloads the public base model. It does not upload your examples.
python -m venv .venv
# macOS / Linux
source .venv/bin/activate
# Windows PowerShell: .venv\Scripts\Activate.ps1
# Install PyTorch for your hardware first; see README.
pip install -r requirements.txt
python validate_data.py
python compare.py --split validation --output baseline-validation.jsonThe 72 fictional examples are a learning scaffold. They are not a production benchmark. Review your model license, data permissions, hardware, and dependency versions before a real run.
The full script checks your files, converts each example into a prompt and target completion, checks lengths, then calls the trainer. The excerpt below shows the key settings; save the full starter for the complete implementation.[2][13]
adapter = LoraConfig(
r=8, lora_alpha=16,
target_modules=["q_proj", "v_proj"],
lora_dropout=0.05, bias="none",
task_type="CAUSAL_LM",
)
settings = SFTConfig(
output_dir="juniper-checkpoints",
num_train_epochs=3,
learning_rate=1e-4,
per_device_train_batch_size=1,
gradient_accumulation_steps=4,
max_length=512,
completion_only_loss=True,
eval_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="eval_loss",
greater_is_better=False,
report_to="none",
)
trainer = SFTTrainer(
model=model, processing_class=tokenizer,
args=settings, peft_config=adapter,
train_dataset=train_data,
eval_dataset=validation_data,
)
trainer.train()
trainer.save_model("juniper-adapter")python train_lora.py
# CPU smoke test, not a useful full training run:
python train_lora.py --cpu --max-steps 2Example settings, not optimal values. The full script adds hardware-aware precision, a fixed seed, length checks, and checkpoint limits. It does not use the browser toy’s learning rate.
Use validation for iteration. Once the configuration is frozen, run the final test on both versions. The comparator counts exact-schema validity, correct routing, and per-class results, and saves every generated reply for inspection.
# Use validation while you are still changing things.
python compare.py --adapter juniper-adapter --split validation
# Only when the configuration is frozen:
python compare.py --base-of juniper-adapter --split test \
--output baseline-test.json
python compare.py --adapter juniper-adapter --split test \
--output tuned-test.jsonTwelve test messages are too few to support a broad reliability claim. Add a representative, independently reviewed evaluation set before release. A higher score does not prove permission, security, or safe business actions.
The adapter folder is not a standalone LLM. The inference script reloads the recorded base model revision and attaches the learned adapter. It validates the JSON before displaying it, and marks invalid outputs for human review.[13]
python predict.py --adapter juniper-adapter \
--message "My mug arrived shattered."
# Check the starting model with the same input:
python predict.py --base-of juniper-adapter \
--message "My mug arrived shattered."Deploy behind your application’s validation and access controls. Fine-tuning does not authorize refunds, grant database access, or provide automatic updates when a policy changes. Keep a version record and a route back to the last known-good model.
Starter status: syntax and dataset checks passed. The Hugging Face/Qwen training run was not executed here because dependencies could not be downloaded. No trained adapter or LLM benchmark is included; see README.md.
The workflow is similar: choose a currently supported model, upload the required data format, create the job, inspect its status and validation metrics, then use the resulting model identifier. The provider operates the training infrastructure; your team still owns the examples, evaluation, permissions, and release decision.
Exact schemas, availability, minimum dataset sizes, and billing vary. Use the chosen service’s current documentation rather than assuming any model or subscription supports fine-tuning. This page does not create a paid job or ask for API keys.[14]
A saved model is an artifact. A trustworthy product is a bigger job.
Ticking these boxes records your review; it does not verify or certify readiness.
Don’t ask whether it learned
the examples.
Ask whether it learned the job.
Not by itself. A PDF can be stored, searched, or included in a prompt without training. Fine-tuning requires a training process that changes trainable parameters. For our SFT project, convert the desired task into reviewed input/output demonstrations; don’t assume a document upload automatically does that.
There is no universal number that makes a model ready. Coverage, label quality, task difficulty, and the starting model matter. The eight browser examples and 72 starter examples exist to teach the workflow, not establish a minimum. Start with a reviewed set, compare against a baseline, and add examples for measured failure patterns.[3]
No such guarantee follows from fine-tuning. Keep changing facts in a current source or authorized service. Test unsupported claims and outdated instructions. The source article identifies overfitting and forgetting as risks; they belong in the evaluation plan, not just the footnotes.
Not in the workflow shown here. At inference time, the model uses the saved parameters to make predictions. A chat history can influence a response without updating weights. A system that collects feedback for later training needs a separate, explicit data and training pipeline.[5]
The supplied article discusses those domain examples, but it does not establish that a fine-tune is safe or reliable for diagnosis or legal advice. Specialized vocabulary is not professional judgment. This guide deliberately uses fictional shop routing; high-stakes applications need domain-specific validation, oversight, and requirements beyond this tutorial.
Keep the goal small, the examples good, and the final exam separate.
Expanded from the supplied AI IXX fine-tuning article: its definition, chef metaphor, dataset categories, preparation → selection → configuration → training → evaluation → deployment sequence, and expert Q&A provide the backbone. The labs, practical starter, prompt/RAG comparison, split guidance, and adapter explanations are additions grounded in the primary sources below. These additions also qualify the original article’s stronger claims about guaranteed improvement and domain expertise.
The browser classifier trains a 4-class linear softmax model. Loss, predictions, gradients, and plots are calculated, not scripted improvements. JSONL checks, LoRA parameter counts, and budget arithmetic run locally.
The hero, story, routing situations, split cards, review checklist, and quiz are authored teaching examples. Juniper is fictional. No provider prices or production benchmark results are claimed.
The Python starter is a separate LLM fine-tuning project, not the browser classifier. Running it downloads model files and uses your compute. There is no paid job, live LLM call, account connection, or key entry in this article.
Documentation checked September 18, 2026. APIs and model availability can change. Implementation status and test limits are recorded in the starter README.