What Is Few-Shot Prompting? How to Use It and How Many Examples to Give

What is few-shot prompting?
Few-shot prompting is a technique where you place a handful of solved examples inside the prompt when you describe a task to an AI model. The model reads the input and output pattern in those examples and then answers a new input in the same way. You do not retrain the model; the learning happens only inside that prompt.
I have worked on digital marketing projects since 2012, and in recent years my team and I have brought AI into our daily workflow. The clearest lesson I have learned is this: showing two or three good examples usually produces far more consistent results than writing long instructions. In this article I explain few-shot prompting not only in theory but also how we use it for marketing and SEO work.
That said, I will not go deep into general prompt writing or chain of thought here. My focus stays on one thing: teaching a model a task by showing it examples.
Where does the idea of few-shot prompting come from?
The concept reached a wide audience through the 2020 paper "Language Models are Few-Shot Learners". Specifically, OpenAI researchers introduced GPT-3 and showed that the model could adapt to new tasks from a few examples in the prompt, without any weight updates.
First, the paper compares three settings. In the zero-shot setting, the model sees only the task description. Next, in the one-shot setting, the researchers add a single example. Finally, in the few-shot setting, they add as many examples as fit into the model's context window. In short, the authors call this ability in-context learning.
The practical value of this finding runs deep. Instead of training a separate model for every new task, you can steer the same model to different jobs by choosing the right examples. That is essentially what you do today when you work with ChatGPT, Gemini or Claude.
Another key finding of the paper: the ability grew stronger as models grew larger. Small models gained little from examples, while large models gained much more. As a result, few-shot methods came to the fore in the era of large language models.
What is the difference between zero-shot, one-shot and few-shot?
At first glance, the difference between the three approaches looks like a simple count of examples. In practice, however, each one shines in different situations. The table below, for instance, sums up the difference.
| Method | Examples in prompt | Where it works well | Weakness |
|---|---|---|---|
| Zero-shot | None | General, familiar tasks; quick tests | Format and tone can vary |
| One-shot | 1 example | Showing a simple format | The model may copy the single example too closely |
| Few-shot | Usually 2 to 5 examples | Custom formats, brand tone, classification | Longer prompt, example selection takes effort |
In short, zero-shot means telling the model what to do, while few-shot means showing it how to do it. The more specific the task is to your business, the more value examples add.
How does few-shot prompting work?
In practice, the model reads the examples in the prompt as a pattern. Specifically, it infers what kind of input arrives, how long and in which format the output should be and which labels apply. Then it applies the same pattern to the final input.
There is also an interesting research finding here. A 2022 study by Min and colleagues showed that randomly replacing the labels in the demonstrations hurt performance on several classification tasks far less than expected. According to the researchers, the label space, the distribution of input text and the overall format drive most of the effect.
I would not read this as "correct examples do not matter". My takeaway differs: the model mainly picks up format and domain from examples. Therefore your examples need a consistent format and a topic close to your real work. You still have to check accuracy yourself, because examples do not guarantee truth.
When should you use few-shot prompting?
First, not every task needs examples. To answer a simple question or summarise a general text, zero-shot usually works well enough. Few-shot prompting makes a difference in these situations:
- When you need a custom output format: a table, JSON or a specific heading structure.
- When you want to keep a brand tone: friendly, corporate or playful?
- When you classify: for example sorting reviews into positive, negative and neutral.
- For short, constrained copy: ad headlines, meta descriptions, short product blurbs.
- When a style resists description: show it rather than saying "write like us".
On the other hand, if the model already performs well on the task, adding examples only lengthens the prompt and raises costs. So try zero-shot first, and add examples if the result falls short.
How many examples should you give?
There is no fixed number, but the official guides suggest a range. Anthropic's Claude prompting best practices recommend 3 to 5 examples for best results. Google's Gemini prompting strategies page recommends always including few-shot examples, but it also warns that too many examples can make the model overfit its responses to them.
My own rule of thumb stays simple. First, I start with three examples. If the output still varies, I look at the type of error and add one example that addresses it. When I pass five, I usually realise the problem lies in example quality, not quantity.
- Simple format task: 2 to 3 examples.
- Classification: at least one example per class.
- Brand tone: 3 to 5 examples on different topics.
Before adding more examples, ask yourself what the model gets wrong. If the format breaks, simplify the format of your examples. If the tone drifts, choose more characteristic examples. Finally, if the content goes wrong, the problem probably lies in the information you gave the model, not in the examples. In other words, that distinction saves you from bloating the prompt for nothing.
How do you choose a good few-shot example?
In practice, example selection forms the most critical step of the technique. Anthropic's guide recommends three qualities. First, examples should be relevant, meaning they mirror your actual use case. Second, they should be diverse; they need to cover edge cases and vary enough that the model does not pick up unintended patterns. Third, they need a clear structure that keeps them apart from the instructions.
- Pick real data: use successful work from your own archive instead of made up samples.
- Ensure variety: if every example comes from one product category, the model struggles with the others.
- Add an edge case: a very short input, missing information or an unusual request.
- Keep quality high: the model copies the mistakes in your examples too.
- Match the target length: long examples lead to long outputs.
For headline work, you can score candidates with the headline analyzer before you pick examples. As a result, the model only sees strong samples.
How should you format examples in the prompt?
Consistent formatting acts as the quiet hero of few-shot prompting. Google's guide stresses that examples should share the same structure and format, otherwise responses may come back in unwanted formats. Anthropic, for its part, suggests wrapping each example in <example> tags and multiple examples in <examples> tags.
In practice I use this template:
- First, describe the task in one or two sentences.
- Then list the examples with the same labels, such as "Input:" and "Output:".
- Use the same fields in the same order in every example.
- Finally, add the new input with the same label and leave the "Output:" line empty.
Put simply, this structure shows the model clearly where the examples end and where its own answer starts. Moreover, a colleague who reads the prompt later can easily tell instructions apart from examples.
Does the order of examples affect the result?
Yes, it can. The "Calibrate Before Use" study by Zhao and colleagues showed that few-shot performance can swing widely depending on the choice and order of examples. Specifically, the researchers identified three biases.
- Majority label bias: the model tends to favour the label that appears most often in the examples.
- Recency bias: the model tends to repeat the label of the last example in the prompt.
- Common token bias: the model tends to prefer labels that appear often in its training data.
The researchers worked with older models, and current models behave in a more balanced way. Still, the practical lesson holds. In classification prompts, give a balanced number of examples per class and mix their order instead of grouping them by class. That way you reduce the tendency to copy the last example.
How do you build a few-shot prompting example for marketing?
Let me show a concrete case. For example, imagine we want short product descriptions for an e-commerce client. With a zero-shot prompt, the model writes generic, flowery texts that all sound alike. With a few-shot prompt, we give three real descriptions from the brand as examples.
- Task: "Write a product description of at most 40 words in the format below. Avoid exaggerated adjectives."
- Example 1: product name and features, followed by the brand's live description.
- Example 2: a product from another category, in the same format.
- Example 3: a tricky product with few features, in the same format.
- New input: the name and features of the target product, with an empty output line.
Order matters here as well. We do not put the tricky product last, because the model may produce outputs that resemble the final example too closely. Also, every example input contains the same fields: product name, material, use case and key feature.
With this setup, the model picks up sentence length, adjective use and the order of information from the examples. As a result, the team edits instead of writing from scratch. Still, we check product facts before anything goes live; the model learns format from examples, not truth.
How do you use few-shot prompting for SEO tasks?
In addition, SEO involves many repetitive tasks with a clear format. Therefore few-shot prompting works very efficiently here. These are the areas my team uses most:
- Meta description drafts: examples that respect the character limit and use the focus keyword naturally. We check drafts in the Google SERP preview tool.
- Keyword intent classification: one example each for informational, commercial, transactional and navigational intent. We build the list first with the keyword suggestion tool.
- FAQ drafts: we show the question plus a 40 to 70 word answer format.
- Schema markup: we show the JSON-LD format, take a draft and then validate it with the schema generator.
I covered how AI search changes SEO in is SEO dead? running SEO and GEO together.
Does the few-shot method work for ad copy?
Yes, especially for ad formats with character limits. In Google Ads responsive search ads, headlines and descriptions must stay within tight character limits. With a zero-shot prompt, the model often exceeds these limits or repeats the same pattern.
In a few-shot prompt, we give headlines that performed well in the past as examples. Then, next to each one, we note which message type it carries: benefit, price, trust or urgency. That way the model learns both length and message variety from the examples.
One warning, though: variations the model generates only count as good once a test proves it. In our Google Ads management work, my team and I run these drafts through human review, check them against ad policies and then test them with real data.
How do you keep a brand voice on social media with few-shot?
First, a brand voice resists description. Phrases like "friendly but professional" mean different things to every model and every person. Few-shot prompting offers the most practical fix, because you show real posts instead of describing them.
Our method works like this: we pick the brand's five posts with the highest engagement. Also, we make sure they cover different topics, such as a campaign, a tip, the team and a customer story. Then we give the new topic in the same format. The model picks up sentence length, emoji use and form of address from the examples.
For short copy experiments, the Instagram bio generator on the site can give you starting ideas. If you need a steady content flow, our social media management service sets up this process end to end.
What are the most common few-shot prompting mistakes?
When I review teams' few-shot prompts, I see the same mistakes again and again. However, most of them need only a small fix.
- Near duplicate examples: the model repeats a single pattern and every output starts the same way.
- Inconsistent format: bullets in one example and paragraphs in another leave the model unsure.
- Flawed examples: typos or wrong facts in an example carry over into the output.
- Unbalanced classes: four positive review examples and one negative push the model toward positive.
- Mixing instructions and examples: without labels, the model may treat an example as an instruction.
- Private data: real customer names and contact details in examples create privacy risks under GDPR or KVKK.
The easiest way to spot these mistakes: ask a colleague to read the prompt. If that person cannot tell at once which part is instruction and which part is example, the model probably faces the same difficulty. Also, if outputs keep repeating a pattern, the source usually sits in the examples.
The last item deserves special attention. Choose examples from real work, but always anonymise personal data.
What is the difference between few-shot prompting and fine-tuning?
Both teach a model through examples, but the methods differ completely. With few-shot prompting, examples travel inside every prompt and the model does not change permanently. With fine-tuning, you run additional training on the model itself with hundreds or thousands of examples.
- Speed: few-shot works instantly; fine-tuning needs data preparation and training time.
- Cost: few-shot uses more tokens per request; fine-tuning needs an upfront investment.
- Flexibility: you can swap few-shot examples within minutes.
- Scale: for a very specific task that runs thousands of times a day, fine-tuning can make sense.
Maintenance also differs. When your brand language changes, you need to retrain a fine-tuned model, whereas with few-shot you just replace the examples. In addition, fine-tuning only works with models that the provider supports for it.
For most marketing teams, few-shot prompting suffices and stays far more practical. I suggest you consider fine-tuning only when you truly hit the limits of few-shot.
Can you combine few-shot with chain of thought?
Yes, you can. If your examples show not only the result but also a short rationale that leads to it, the model follows a similar reasoning style. This approach can help especially in classification tasks that require judgement.
For example, when you label keyword intent, you can add a "rationale" line to each example: "The keyword contains a price, the user sits close to purchase, therefore commercial." For a new keyword, the model then writes the rationale first and the label second.
That said, chain of thought forms a broad topic of its own and deserves a separate article. Here I only want to stress one point: adding rationales makes the prompt longer. So use this only for tasks that truly require reasoning, where simple examples fall short.
How do you test and measure your few-shot prompts?
You judge a good prompt by measurement, not by gut feeling. That is why I recommend testing few-shot prompts with a small test set. The process stays simple.
- Pick 10 to 20 inputs from your real work that you did not use as examples.
- Write down your criteria for a good output: length, format, tone, accuracy.
- Run a zero-shot prompt first, then the few-shot prompt.
- Score the outputs against the criteria; ideally two people score separately.
- Note the error types and add examples that target them.
- Repeat the test whenever the model version changes.
Many teams skip the last step. Yet model updates can change how a prompt that worked perfectly yesterday behaves. If you keep your test set, you notice that change within minutes.
Does few-shot prompting give the same result on every model?
No, it does not. Models differ in training data, size and training method. Therefore the same few-shot prompt can produce different results on ChatGPT, Gemini and Claude. Current models also follow instructions better than older ones and need fewer examples for some tasks.
My practical advice: read the official guide of the model you plan to use. The guides from Anthropic, Google and OpenAI all cover the use of examples, but their advice on tag format and example count differs slightly. If you want to understand how these models work in more depth, take a look at my article on large language models (LLMs).
Also, whenever you move a prompt from one model to another, rerun your test set. An example set that works perfectly on one model can feel too long or too thin on another.
Is there a ready made template for few-shot prompting?
You will need to adapt it for each task, but the skeleton below offers a good start for most marketing jobs. Fill the brackets with your own content.
- Role and task: "You are an editor who writes [task] for [brand]."
- Rules: length limit, banned phrases, required elements.
- Example block: 3 to 5 examples, each with "Input:" and "Output:" labels, inside an <examples> tag.
- New input: with the same "Input:" label.
- Output slot: an empty "Output:" line.
I suggest you keep this template in a shared team document. That way everyone uses the same structure, nobody loses good example sets and new colleagues adapt quickly.
When does few-shot prompting fall short?
Few-shot prompting offers a powerful technique; however, it does not solve every problem. In three situations you hit its limits quickly.
- Missing knowledge: if the model does not know a topic, examples do not give it that knowledge. For current prices, regulations or stock levels, you need to hand the source to the model directly.
- Multi step reasoning: for tasks that need calculation or a chain of logic, examples that show only the result often fall short.
- Very long outputs: for long reports or articles, examples inflate the prompt too much.
In these cases the fix usually does not involve more examples. Instead, you can add the relevant document, split the task into small steps or write a separate prompt for each step.
My own rule: if I still see the same error after two rounds of example tweaks, I accept that the problem lies in task design, not in the examples. Then I redefine the task. This saves far more time than swapping examples for hours.
Should you include negative examples?
Some teams add bad examples to their prompts with a note like "do not write like this". Used carefully, this can help; however, it also carries risk. The model tends to treat every text it sees in the prompt as a pattern. As a result, the words and structure of a bad example can leak into the output without you noticing.
In my experience, the safer route states prohibitions as rules and shows only the correct behaviour in examples. For instance, writing the rule "avoid exaggerated adjectives" and keeping all examples free of them yields more consistent results.
- Start with positive examples only.
- If a specific error persists, state it clearly as a rule.
- If the error still persists, add the negative example with a clear label such as "Wrong example:".
- Right after the negative example, show its corrected version.
This sequence focuses on teaching the model what right looks like, not what wrong looks like.
How do you build a few-shot example library for your team?
Good examples form a team's most valuable AI asset. Yet in most teams they vanish in personal chat histories. For this reason I recommend a shared example library.
- Organise by task: product descriptions, ad headlines, FAQs, meta descriptions and so on.
- Add a source to each example: note which campaign it came from and how it performed.
- Add a date: prune old examples when your brand language changes.
- Remove personal data: only anonymous examples go into the library.
- Assign an owner: one person keeps the library up to date.
Over time this library turns into your brand's written memory. When a new colleague joins, they see what good examples look like before they learn to write prompts. Moreover, content from different people ends up in the same tone.
How do you balance cost and length in few-shot prompts?
Every example lengthens the request and therefore raises token usage. For one off jobs the difference stays negligible. However, if you run the same prompt hundreds of times a day through an API, the number of examples shows up directly on your bill.
I use a few practical methods to strike a balance. First, I trim examples and remove every sentence the task does not need. Second, I use the test set to measure how many examples really make a difference. Often the fifth example brings no measurable gain over the third.
- Strip examples of unnecessary explanations.
- Keep the fixed part of the prompt identical across calls; some providers offer caching for repeated prompt sections, so check your provider's documentation for details.
- Route simple tasks to smaller, cheaper models.
In short, the goal is not to give the most examples but to find the fewest examples that reach the quality you need.
Where should you start learning few-shot prompting?
The best way to learn: run a test with one task from your own work. Pick a short writing job that costs you time every week. Try zero-shot first, then add three good examples from your archive and compare. This small experiment teaches more than dozens of articles.
Next, read the official documentation. The Anthropic and Google guides I linked in this article stay short and practical. Then build your test set, so you can improve your prompts over time.
If you are thinking about how AI fits into your overall marketing strategy, my team and I address both content production and visibility in AI search within our SEO consulting work. To sum up, few-shot prompting ranks among the techniques that deliver the fastest quality gain at the lowest cost. Choosing the right examples, however, depends entirely on your expertise.




