The Teacher Sets the Ceiling: What Distilling an LLM Classifier Into a Small Model Actually Buys You
An experiment replacing a mid-size generalist LLM classifier with a fine-tuned small model. Model size, data volume and constrained decoding made no measurable difference. Changing who wrote the training labels made the largest one. And without ground truth, every metric turned out to measure agreement with a judge, not correctness.
Porsync Research · Published 2026-09-28
Finding
When distilling an LLM classifier, the label source sets the accuracy ceiling: a small student matches its teacher regardless of size, data volume or grammar, while retraining the identical recipe on a stronger teacher's labels moves the result by about ten points of balanced accuracy. Without outcome data, that improvement is agreement with a better-specified judge, not proof of better decisions.
Small-model distillation experiment — findings (Sep 2026)
A small LoRA student versus the larger generalist model whose labels it was trained on, plus ablations and a teacher swap, on a three-class scoring task with a binary accept/reject decision. Local hardware, one consumer GPU.
| Measurement | Value | Note |
|---|---|---|
| Student vs teacher accuracy | Tie | within the teacher's own run-to-run spread |
| Throughput | ~12× faster | small student vs the generalist teacher |
| Run-to-run stability | Deterministic | the student's decisions never changed across seeded repeats; the teacher's did |
| Smaller base · less data · no grammar | No measurable effect | each ablation indistinguishable from the reference |
| Teacher swap (same recipe) | ≈ +10 pts balanced accuracy | statistically significant on the accept/reject decision |
| Ground truth available | None | every metric is agreement with some labeller |
The question
A classifier inside one of my automation pipelines runs on a mid-size generalist LLM. It works, but it's slow and its outputs drift between identical runs. The question was the obvious one: can a small fine-tuned model — a LoRA on a ~2B base — replace it?
The pass/fail thresholds were written down before any training ran, and every comparison used a fixed held-out set labelled against a written rubric.
Finding 1: the student matches the teacher, and nothing about the student matters
Trained on the generalist's own labels, the small model tied it on the held-out set, ran roughly **12× faster**, and was **deterministic** — its decisions never moved across seeded repeats, while the teacher's did.
Then the ablations, one variable at a time: a smaller base model, a quarter of the training data, a few percent of the training data, and constrained decoding switched off. **None of them made a measurable difference.** A student trained on a small fraction of the labels scored the same as one trained on all of them.
That pattern has one explanation. A student can only learn its teacher's judgement. When size, data volume and decoding all stop mattering, the constraint isn't the student — it's the labels.
Finding 2: change the teacher, and the ceiling moves
So I kept the recipe and changed the teacher: a frontier model re-labelled the same data against the same written rubric, and I retrained the identical small model on those labels. (One trap worth naming: the new split had quietly pulled some held-out items into training. Removing them first is the difference between a result and a leak.)
On the accept/reject decision, balanced accuracy rose by roughly **ten points**, significantly beating both the original student and the generalist it was meant to replace. Same model, same hyperparameters; only the label source changed.
Finding 3: a bad metric was really a bad answer key
The three-class result got *worse*, and the loss was one class the new model never predicted. Tracing it back found two answer keys that disagreed about what that class meant — and one of them was partly noise, because a batch of held-out labels had been assigned by a category rule with a random pick inside a score range. No model can predict a coin flip.
Making the rule explicit and applying it to both the training labels and the held-out set fixed it: the rare class went from never predicted to the model's most precise call.
Finding 4: without ground truth, you are choosing a judge
Last, the retrained model ran in shadow against production on live traffic, writing to a separate table so production was never touched. It disagreed with the incumbent far more than the original student did — which is exactly what retraining on a different teacher should produce.
This is where the study has to stop claiming things. **There was no ground truth.** The held-out set is one rubric. The new teacher is another model's reading of that rubric. The shadow run is agreement with the incumbent. Every number measures agreement with *some* judge — none measures whether a decision was right.
What I'd take from it
1. **Distillation copies judgement; it doesn't improve it.** If the student ties the teacher regardless of size or data, stop tuning the student and look at the labels.
2. **Audit the answer key before blaming the model.** The worst result in this experiment came from the evaluation set, not the model.
3. **A small specialist is an operational win before it is an accuracy win.** Faster, deterministic, and in a fixed memory budget — while matching the incumbent.
4. **Without outcomes, picking a model is picking a judge.** The real test is to route a sample of the cases where two models disagree through the normal process and record what actually happens.
Nothing from this experiment is deployed. Its job was to produce the evidence for that decision, not to make it.