Prompting7 min read
Few-Shot Prompting: When and How to Use Examples
What few-shot prompting is, when examples beat instructions, how many to use, how to pick and format them, and how to stop the model copying them too closely.
Published
Few-shot prompting means adding a handful of worked examples to your prompt, each showing an input and the output you want, so the model can infer the pattern. It is most useful when the result is easier to show than to describe: a particular tone of voice, a classification scheme with subtle boundaries, a fixed output format or a way of handling tricky cases. Two to five varied, high-quality examples are usually enough. Combine them with clear instructions, mark them clearly as examples, and test whether they actually improve results on your own cases.
This article is part of our prompting cluster. For the overall method, start with the prompt engineering guide for business.
Zero-shot, one-shot and few-shot
| Approach | What the prompt contains | Typical use |
|---|---|---|
| Zero-shot | Instructions only | Common tasks the model understands well, such as summarising or translating |
| One-shot | Instructions plus one example | Showing a format or style once |
| Few-shot | Instructions plus several examples | Subtle distinctions, consistent style, structured outputs |
Modern models handle many tasks well zero-shot. Few-shot prompting is not a default to apply everywhere; it is a tool for cases where instructions alone produce inconsistent results.
When examples help most
- House style and tone. Describing "warm but not chatty, direct but not curt" is hard. Two sample replies make it concrete.
- Classification with fuzzy boundaries. If "complaint" and "feedback" overlap, examples of borderline cases show where you draw the line.
- Exact formats. A sample of the JSON, table or report layout you need reduces formatting errors.
- Edge-case handling. An example showing how to respond when information is missing teaches the model to do the same.
- Domain conventions. Industry-specific phrasing, abbreviations or structures.
When examples are not worth it
- Simple, well-understood tasks. Adding examples costs tokens without improving results.
- When you cannot produce good examples. Weak examples teach weak output.
- Highly varied tasks. If every request is different, examples may push the model towards the wrong pattern.
- Very long examples. If each example is a full document, the prompt becomes expensive and the model may focus on the examples instead of the actual input.
A worked example: classifying customer messages
Instructions only (zero-shot):
Classify the customer message as one of: order_status, return_request, complaint,
product_question, other. Reply with the label only.
This works for clear cases but may struggle with messages like "The chair arrived scratched, can I send it back?" Is that a complaint or a return request?
Few-shot version:
Classify the customer message as one of: order_status, return_request, complaint,
product_question, other. Reply with the label only.
Rules:
- If the customer wants to send something back, use return_request, even if they
are also unhappy.
- Use complaint when the customer expresses dissatisfaction without requesting a
return or exchange.
Examples:
Message: "Where is my order? It was supposed to arrive yesterday."
Label: order_status
Message: "The chair arrived scratched, can I send it back?"
Label: return_request
Message: "Your delivery driver was rude and left the box in the rain."
Label: complaint
Message: "Is the oak desk also available in 160 cm?"
Label: product_question
Now classify:
Message: """[customer message]"""
Label:
Notice three things: the rules state the decision explicitly, the examples cover the tricky boundary, and each label appears once. The examples support the rules; they do not replace them.
How many examples?
There is no universal number. A practical approach:
- Start with zero-shot and a clear instruction.
- If results are inconsistent, add two or three examples covering the problem cases.
- Test on a fixed set of real inputs.
- Add more examples only if they fix specific failures.
Remember that in an application, examples are part of the input on every request. Five examples of 150 tokens each add 750 tokens per call. Over many requests that adds up; the LLM API pricing guide shows how to calculate it, and prompt caching can reduce the cost of a stable example block.
Choosing good examples
- Representative: examples should look like real inputs, not idealised ones.
- Varied: different lengths, phrasings and situations, so the model learns the pattern, not the surface.
- Balanced: in classification, avoid having most examples of one label; the model may favour it.
- Correct: every example output must be exactly what you want. A single sloppy example can be copied.
- Edge cases included: at least one example of the situation where the model tends to fail.
- Short: trim examples to what demonstrates the point.
Formatting examples
Clear structure helps the model separate examples from the real task:
- Use consistent labels such as "Message:" and "Label:", or "Input:" and "Output:".
- Separate examples with blank lines or delimiters.
- Optionally wrap them in tags, for example
<example>...</example>. - Put the real input last, in the same format as the examples, followed by the output label.
- State explicitly: "The examples illustrate format and style. Do not reuse their content."
Avoiding over-copying
A common problem is that the model mirrors examples too closely, reusing phrases or forcing every answer into the same structure. To reduce it:
- Vary example wording and length.
- Use examples from different contexts.
- Say what should transfer ("tone and structure") and what should not ("specific wording, names, facts").
- Avoid examples that contain specific facts the model might repeat in unrelated answers.
- If over-copying persists, use fewer examples or describe the pattern in words instead.
Few-shot for writing style
For drafting tasks, examples of good past writing are often the most efficient way to convey style. A pattern that works well:
Write a reply to the customer email below in our house style.
Our style: friendly, direct, short sentences, no exclamation marks, always end with
one clear next step.
Two examples of replies we consider good (for style only, not content):
<example>[past reply 1]</example>
<example>[past reply 2]</example>
Customer email:
"""[email]"""
Remove personal data from past replies before using them as examples, in line with your company's AI policy.
Few-shot in system prompts
In applications, examples often live in the system prompt so they apply to every conversation. Keep them short, clearly marked, and review them whenever you update the rules, so examples and instructions never contradict each other. How to write a system prompt covers where examples fit in the overall structure.
Few-shot prompting versus fine-tuning
Both teach a model through examples, but very differently:
| Few-shot prompting | Fine-tuning | |
|---|---|---|
| Where examples live | In the prompt, every request | Built into the model through training |
| Number of examples | A handful | Typically many more |
| Setup effort | Minutes | Data preparation, training, evaluation |
| Cost structure | More input tokens per request | Training cost plus model usage |
| Flexibility | Change examples any time | Retraining needed to change |
Start with few-shot prompting. Consider fine-tuning only when you have many high-quality examples, high volume and a stable task. See RAG vs. fine-tuning for a fuller comparison.
Testing whether examples help
- Build a test set of real inputs, including difficult ones.
- Run the prompt without examples and record results.
- Run with examples and compare.
- Look at failures in both versions. Did examples fix them, or introduce new ones?
- Keep only the examples that earn their place.
Common mistakes
- Examples that contradict the instructions. The model may follow the example rather than the rule.
- All examples alike. The model learns the narrow pattern.
- Unbalanced labels. The model leans towards the most frequent one.
- Examples with errors. They get reproduced.
- Using examples instead of instructions. Combine both; state the rule, then show it.
- Never revisiting examples. Update them when your policy, products or style change.
Try it
The AI prompt generator has a field for an example of a good result and adds a note telling the model to follow its style and structure rather than its content. Paste one of your best past outputs there and compare the result with the version without an example. For more starting points, see prompt templates for business.
FAQ
What is few-shot prompting?
Few-shot prompting means including a small number of examples of input and desired output in a prompt, so the model can infer the pattern, style or format you want. Zero-shot means no examples; one-shot means a single example.
How many examples should a few-shot prompt contain?
Often two to five are enough. Start with a few varied examples and add more only if testing shows a benefit. More examples cost more tokens on every request and can make the model over-fit to them.
When is few-shot prompting better than instructions?
When the desired result is easier to show than to describe, such as a house style, a classification scheme with subtle boundaries, or a precise output format.
Why does the model copy my examples too closely?
If examples are similar to each other or very specific, the model may reuse their wording or structure. Vary the examples, say they illustrate the pattern rather than the content, and keep instructions explicit.
Related articles
Prompting8 min read
Prompt Engineering for Business: A Practical Guide
Prompt engineering for business users: a six-part prompt structure, before-and-after examples, how to test prompts and how to share them in a team.
Prompting7 min read
How to Write a System Prompt (With Examples)
How to write a system prompt for an AI assistant or automation: a proven structure, a full example for customer support, testing tips and the mistakes to avoid.
Prompting7 min read
Prompt Templates for Everyday Business Tasks
Twelve ready-to-use prompt templates for business: emails, summaries, meeting notes, data extraction, reports and feedback, with tips to adapt them.