What Is Knowledge Distillation? How a Small AI Model Learns From a Big One

What is knowledge distillation?
Knowledge distillation, also called model distillation, is a technique that teaches a smaller AI model (the student) to imitate a larger, stronger model (the teacher). The student copies the teacher's probability estimates, not only its final answers. As a result, it gets close to the teacher's quality while using far fewer resources.
Think of a master chef and an apprentice. The chef does not write down every recipe. Instead, the apprentice works next to the chef and absorbs the timing, the taste, and the hand movements. However, the apprentice will not match the chef's full experience. Still, they can cook a narrow menu much faster and cheaper.
First, a note on scope: this article covers one term only. We mention neighboring concepts in the comparison table and in short asides, and we link to separate guides for them.
How does the teacher and student analogy help?
Let us stretch the analogy a little. The teacher model is large, expensive, and slow. It may have billions of parameters, but we leave exact figures out on purpose, because current values change and you should check the provider's official documentation. The student has to do the same job with much less computation.
A good classroom teacher never says only "the answer is B." Instead, the teacher says "I think B, C is plausible, and A is clearly wrong." That extra shading is what the student really gains.
The analogy has limits, though. A student cannot learn a topic the teacher never covered. It also inherits the teacher's mistakes very easily. In short, distillation is not magic compression. It is a deliberate transfer of knowledge.
Also, the word itself comes from chemistry. There, distillation separates the essence of a mixture into a more concentrated form. Here, the mixture is the large model's knowledge, and the essence is the behavior you move into the small model.
What does the original distillation paper say?
The best-known source is the paper "Distilling the Knowledge in a Neural Network" by Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. It starts from a practical problem: running a whole ensemble of models for every prediction is cumbersome and can be too expensive. The proposed fix is to compress the ensemble's knowledge into a single model that is easier to deploy.
The abstract points to three main ideas:
- The student learns from the teacher's softened probabilities as well as from hard labels.
- A parameter called temperature controls how soft those probabilities become.
- Specialist models can separate fine-grained classes that the general models confuse.
For the details, you can read the full text on the arXiv page. Here we only carry over the concept, and we do not copy numbers from the experiments.
Then the idea moved into language models. The Hugging Face team shared a distilled version of the BERT language model in the DistilBERT paper. It showed that distillation also works for language understanding, not only for classification. Read the source for the exact results.
How does knowledge distillation work step by step?
The process has four conceptual steps. We keep each step simple, because details vary from one framework to another.
- First, you pick or train a strong teacher model.
- Next, you show the teacher many example inputs and record its outputs.
- Then you train the small student to imitate both the real labels and the teacher's outputs.
- Finally, you test the student on real task data and compare it with the teacher.
The key idea is that the student's target is not just the "right class." It tries to capture the teacher's whole probability distribution. Therefore, the student learns from far more signal than a single answer carries.
The training loss usually has two parts. One part measures agreement with the true label. Meanwhile, the other measures agreement with the teacher. Your team tunes the weight of each part, and the best balance depends on the task.
Also, the "transfer set" matters here. The examples you show the student must represent real use. If they are narrow or biased, the student learns only that narrow world. So build the example pool from real inputs and include rare cases too.
What is a soft label and why does it carry more information than a hard label?
A hard label names the single correct class for an example. For a photo, saying "cat" is a hard label. A soft label is the probability the model assigns to every class: high for cat, medium for dog, very low for truck.
Why does that matter? Because it reveals similarity between classes. When the teacher says "this cat looks a bit like a dog," the student also learns how close those animals are. A hard label erases that information completely.
Here is a small example calculation. Suppose the teacher gives 0.90 to cat, 0.09 to dog, and 0.01 to truck. The hard label says only "cat." The soft label also says that a cat is much closer to a dog than to a truck. These numbers only illustrate the concept and are not real measurements.
In practice, soft labels bring three benefits:
- Generalization improves, even with fewer examples.
- The student recognizes the borderline cases where the teacher hesitates.
- Training of the small model also becomes more stable, because the signal is richer.
The literature sometimes calls this "dark knowledge." The phrase is a metaphor. Nothing mysterious is going on; the signal is simply the fine differences between probabilities.
What does temperature do in distillation?
Temperature is a parameter that controls how sharp or soft a probability distribution is. At low temperature, the model locks onto almost one choice. At high temperature, the probabilities move closer together, and the signal from the second and third choices becomes visible.
For example, think of a photo. In a harsh photo you only see the brightest spot. In a softened photo, the details in the shadows appear too. Temperature works like a dimmer that lights up those shadows, so the student sees how the teacher weighs its alternatives.
The original approach applies this softening during training. After training ends, the temperature returns to normal.
The `temperature` setting you adjust when generating text serves a different purpose: it controls variety in the output. Both rest on the same mathematical idea, but their roles differ. We cover the generation side in our guide to temperature and top-p.
How do you distill a large language model with synthetic outputs?
Large language models change the picture a little. Often you cannot see the teacher's internal probabilities, so the setup changes. You only get text. In that case the common approach is simple: you give the teacher a list of tasks, collect its answers, and use those answers as training data for the small model.
Here the student imitates the teacher's written text. So "synthetic output" takes the place of the soft label. We discuss the general logic of artificial training data in our article on synthetic data, so we do not repeat it here.
Still, the strength of this approach is convenience. The weakness is that the teacher's errors, style, and biases flow into the student. Also, a small model weakens quickly once it leaves the area the teacher covered.
Diversity of the data decides a lot. Prepare tasks from many topics, lengths, and difficulty levels instead of a hundred similar questions. Then remove duplicates. In addition, treat the cases the teacher refused or could not answer as their own class, because the student must know those limits as well.
A good habit is to filter the teacher's outputs before training. For example, drop wrong, contradictory, or needlessly long answers.
What are the main types of distillation?
You can group distillation by what the student imitates. The split is conceptual, and frameworks may use other names.
- Response-based distillation: The student imitates the teacher's final output or probabilities. This is the classic approach in the Hinton paper.
- Intermediate-layer distillation: The student also tries to match representations inside the teacher's layers. This route needs access to the teacher's internals.
- Sequence-level distillation: The student takes the full sentences or answers the teacher wrote as examples. Training on synthetic outputs for language models sits close to this group.
Your choice therefore depends on how much access you have to the teacher. If you only receive text from a closed API, the first route is often unavailable. With an open-weight model, you can consider all three.
The split also touches permissions. Distilling from a model whose weights you own differs, both technically and legally, from distilling from someone else's API output. We return to licensing in a separate section below.
What is knowledge distillation and why did it spread so widely?
The short answer is cost, speed, and location. Large models are powerful, but every request needs heavy computation. As traffic grows, the bill grows and response time stretches. Running a large language model for every small task is often wasteful.
Distillation changes that balance. If you train a small model for a narrow task, you gain:
- Lower latency: answers arrive faster.
- Lower unit cost: the same hardware serves more requests.
- Less memory: the model can run on limited devices such as phones or browsers.
- More control: you can host the model in your own environment.
We give no numeric comparison. The gain differs from task to task, and you need to measure it in each project. If you want to understand token-based pricing, see our guide to tokens and API cost.
Product cycles also play a role. For example, teams first build a quick prototype with a large model. Then, once usage becomes clear, they move the most frequent tasks to small models. The expensive model then runs only on the hard cases that really need it.
Where does knowledge distillation help most?
In practice, use cases appear wherever you say "the big model is too heavy." Here are a few typical ones.
- On-device AI: Keyboard suggestions, offline translation, and voice commands on earbuds all need small models.
- Real-time systems: In live chat, search suggestions, and voice assistants, latency directly shapes the user experience.
- High-volume classification: For tagging support tickets or sorting reviews, thousands of repetitive jobs a day suit a small model.
- Local deployments: Organizations that do not want data to leave their network can run small models on their own servers.
Ask yourself one question when choosing: does this task repeat very often, and is the answer format predictable? If both answers are yes, distillation is a strong candidate. If the answers are open-ended and creative every time, a small model reaches its limit sooner.
What is knowledge distillation in a business scenario?
The following is an example scenario, not a real client result. Imagine an online store. It wants to sort customer messages into "shipping," "returns," "product question," and "complaint."
At first, the company uses a large model for the job. The results are good, but it pays for every message and waits for every answer. As message volume grows, so does the cost.
At this point the team could proceed like this:
- Build a sample pool from real messages after removing personal details.
- Collect the labels the large model gives for those samples.
- Review a portion by hand and fix the wrong labels.
- Train a small model on those labels.
- Compare the small model with the large one on a test set kept apart from training.
If the result is good enough, then the daily load moves to the small model. Unclear messages go to the large model or to a person. This keeps quality and cost in balance. For setups like this, see our AI consulting page.
However, the work does not end at launch. As incoming messages shift, for example when a new campaign brings new questions, you need to watch the small model's confidence. Routing low-confidence answers to the large model is a good safety valve. Those records also become a valuable pool for the next training round.
How do distillation, quantization, and pruning differ?
All three methods lighten a model, but they take different routes, and people mix them up often. Quantization stores the model's numbers with fewer bits. Pruning removes connections that add little. The table below summarizes the differences.
| Method | What it does | Model architecture | Extra training needed? | Typical risk |
|---|---|---|---|---|
| Distillation | Trains a small student on a large teacher's outputs | New, smaller model | Yes, a full training run | Inherits the teacher's mistakes |
| Quantization | Lowers the precision of the numbers | Same model, fewer bits | Usually none or very little | Quality loss on sensitive tasks |
| Pruning | Removes unimportant connections or neurons | Same model, sparse structure | Often a fine-tuning pass | Sudden quality drop if overdone |
| Fine-tuning | Adapts an existing model to a new task | Same size | Yes, but lighter | Can forget earlier knowledge |
Also, you can combine these methods. For example, you distill first and quantize afterward. Then the gains in size and speed can multiply.
Here is a rough rule for choosing. To shrink a ready-made model quickly, quantization takes the least effort. Pruning makes sense when you know many connections are redundant. For the smallest and fastest model on a narrow task, distillation stands out. Because gains vary by task, test each option on your own data.
Is distillation the same as fine-tuning?
No, but they are neighbors. Fine-tuning adapts an existing model to your data, and the model size stays the same. Distillation trains a different, usually smaller model and uses the teacher as the target.
In practice the two often meet. For example, you fine-tune a large model for your domain first. Then you make that expert model the teacher and distill a small student from it. You can find the basics of fine-tuning in our guide to fine-tuning and LoRA.
A simple question helps you decide: is the problem "the model knows the wrong things," or "the model is too heavy"? In the first case, fine-tuning fits better. In the second, distillation does. For missing knowledge, retrieval is often enough.
Another difference is the target signal, and it matters. In fine-tuning, the target is usually a human label. In distillation, the target is the teacher's output. Therefore distillation can work even when you have little labeled data, as long as you can trust the teacher's quality.
In what order should you try these methods?
For most projects, starting with the cheapest path makes sense. The order below is a general approach from field experience, and it may not fit every task exactly.
- First, improve the prompt and add a few examples.
- Then, if knowledge is missing, set up retrieval-augmented generation so the model can read your documents.
- Next, consider fine-tuning if you have a behavior or style problem.
- Evaluate distillation if cost or latency remains a problem.
- Finally, improve size and speed further with quantization.
So this order is a saving tip, not a rule. Distillation is one of the most labor-intensive steps. So do not jump to it before you have measured the problem. Also, measure each step with the same test set. That way you see which method really made a difference.
What are the advantages of distillation?
The advantages show up clearly in the right use case. We can summarize them like this:
- Speed: A small model answers in less time.
- Cost: The same job needs less computation.
- Portability: The model can fit on edge devices and local servers.
- Privacy: You can process data without sending it out.
- Predictability: A model trained for a narrow task can behave more consistently than a general-purpose one.
Moreover, distillation helps a product mature. You build the prototype with a large model, watch the traffic and needs, and then switch to a small model. In other words, you first get a product that works, and then a product that works cheaply.
However, one point deserves emphasis: none of these advantages arrive automatically. All of them depend on a good teacher, clean training data, and a solid test. If one of the three is missing, the small model becomes fast but unreliable.
What are the limits and risks of distillation?
To be honest, distillation is not a free lunch. The main limits are these:
- Capacity limit: A small model cannot carry all of the large model's ability. General knowledge and complex reasoning weaken first.
- Inherited errors: If the teacher is wrong or biased, the student repeats the same flaw.
- Distribution shift: When an input arrives that the student never saw, it can behave unexpectedly.
- Maintenance load: When the teacher changes, you may need to retrain the student.
- Measurement difficulty: Serious errors can hide in rare cases even when the average score looks good.
Therefore, to reduce these risks, prepare a broad test set. Keep human oversight on critical decisions as well. For the hallucination side of the topic, see our article on AI hallucination.
For instance, consider an example scenario. A student trained on billing questions receives a return question one day. Because it never saw that area, it may answer with confidence but incorrectly. So add a filter that recognizes out-of-scope inputs, plus a route to a person.
Can you train a model on another provider's output without license trouble?
This question matters, because distillation in practice often relies on a provider's API output. Many providers may restrict using their outputs to build a competing model in their terms of use. Terms differ between providers and also change over time.
For that reason we suggest these steps:
- Read the current terms of use and service terms of the provider you use.
- Clarify whether using outputs to train another model is allowed.
- Where something stays unclear, ask the provider in writing.
- For open-license models, read the license text separately, because some carry their own restrictions.
- Record your decision and your reasoning.
Opening fake accounts or hiding your method to get around the terms is not a solution. Such attempts count as circumventing the systems and can get your account closed. This section is not legal advice; for critical projects, consult a lawyer.
Open-license models are not simple either, though. Some licenses regulate commercial use, and others regulate training another model on the output. Review the license text each time, and check again when you switch versions.
When is a small model not the right choice?
Not every project suits distillation. In the following cases, staying with the large model may be wiser.
- The task is very broad and must answer any kind of question.
- Traffic is low, so the cost gain does not repay the training effort.
- You have no test set you can trust.
- Nobody on your team can maintain model training.
- The tolerance for error is very low, and every deviation has serious consequences.
In these cases, first try lighter fixes such as better prompts. In short, retrieval often fixes missing knowledge, fine-tuning fixes missing behavior, and distillation fixes cost and speed.
However, there is also a middle path. Leave the heavy work to the large model and hand only the most frequent, easy part to the small one. This hybrid setup removes the need to distill everything. It also shrinks the risk.
What checklist should you follow before a distillation project?
The list below is a practical starting point for software teams and business owners.
- Did you define the task in one sentence?
- Did you measure that the large model performs well enough on this task?
- Do you see the current cost and latency problem in concrete terms?
- Does the training data contain personal or confidential information, and did you remove it?
- Do the teacher provider's terms allow using outputs this way?
- Did you set aside a separate test set and an acceptance threshold?
- Do you have a plan to route cases to the large model or a person when the small model falls short?
- Did you write down who updates the model and how often?
If you cannot answer yes to most of these, finish the groundwork first. So treat the list like a project document: assign an owner to each question and write the answer down. A question without an answer is the riskiest point of the project.
If you need setup help, see our private LLM deployment page and our Ollama guide.
How do you measure the quality of a distilled student?
Measurement is the most neglected step of distillation. Average accuracy alone does not suffice. We recommend looking at three levels.
First, agreement with the teacher: on the same inputs, how closely does the student match? Second, real task success: what is the result on a human-labeled test set? Third, hard cases: how does the model react to rare, ambiguous, or edge cases?
Also, put deliberately tricky examples in your test set. It is not surprising when the student matches the teacher on easy examples. The real gap shows up on borderline ones. So inspect error types one by one, not just the average score.
Next, set up live monitoring in production too. Track user complaints, the share of cases you route to people, and latency. This way you notice a quality drop early.
Evaluation methods change quickly, so we state no specific score or threshold. Define the right threshold for your own task with your team.
What mistakes do teams make most often?
These are the mistakes we see most often in the field:
- Feeding the teacher's output into training without any filtering.
- Mixing the test set with training data and inflating the score.
- Expecting the large model's ability from the small model.
- Collecting outputs without reading the terms of use.
- Skipping monitoring after training.
Each fix is simple, but it takes discipline. Setting the test set aside at the start and never touching it is the cheapest insurance in the project.
Version control is another common gap, because providers change things. If the provider changes the teacher model, the student you trained may age over time. Record which teacher version and which data you used. Then you find the cause quickly when a problem appears.
Conclusion: what is knowledge distillation and where do you start?
Knowledge distillation transfers the knowledge of a large model into a smaller one, with the goal of cutting cost and latency. Through soft labels or synthetic outputs, the student captures the teacher's behavior on a narrow task.
In short, the practical answer to "what is knowledge distillation" is this: the same job done by a smaller, cheaper model. To begin, clarify the task and the current cost. Then read the teacher provider's terms, run a small pilot, and measure the result on a separate test set. Quantization and pruning complement this path, and they do not replace it.
At Talha Aslan and team, we advise choosing the need first and the tool second in AI projects. For the wider picture, see our AI automation services.



