Artificial Intelligence

What Is Prompt Engineering? Techniques, Best Practices and Examples

Talha AslanTalha Aslan 19 min read 2 views

In client meetings I keep hearing the same sentence: "I ask ChatGPT, but what comes back is useless." Most of the time the problem is not the model; it is the request. Prompt engineering closes exactly that gap. In this guide I walk through the core structure, expert techniques, model parameters and before/after examples from marketing work. I also share, honestly, which habits waste the most time.

What is prompt engineering?

Prompt engineering is the practice of designing, testing and refining the instructions you give a large language model so it produces the output you need, consistently. In other words, you turn an unpredictable chat into a measurable workflow by giving the model a role, context, a task, a format and constraints.

The term sounds technical, but at its core it is communication. It resembles a good manager briefing a new colleague. You explain what you want, why you want it and what finished work looks like. The model does not know the context in your head; consequently, it guesses everything you leave out.

The Anthropic prompt engineering documentation also recommends starting with clear success criteria and a way to test against them. I find that approach extremely useful, because a "good prompt" only means something once you have defined a "good result".

Why does prompt engineering matter so much now?

AI tools now sit in every corner of the office. Marketing teams draft ad copy with them, sales teams prepare emails, developers ask for code reviews. However, two people using the same tool get very different quality. In most cases the difference comes from the structure of the request, not from the tool.

In addition, newer models follow instructions more literally. That is good news, but it has a cost: the model also follows a vague instruction literally, and you end up with average text. The OpenAI prompt engineering guide likewise puts clear instructions and sufficient context at the center.

From a business point of view, the issue is efficiency. A weak prompt means five rounds of corrections. A well structured prompt, on the other hand, brings a nearly publishable draft on the first try. In short, prompt engineering directly affects the return on the time you spend with AI.

What are the parts of a good prompt?

Over the years my team and I have settled on a five part structure. You do not need every part in every request; still, using this skeleton for important work improves quality noticeably.

  1. Role: The expertise the model should think with. For example, "You are a copywriter experienced in B2B software marketing."
  2. Context: Audience, brand, product, earlier attempts and hard facts that limit the work.
  3. Task: A one sentence job description. "Write product page headlines from three different angles."
  4. Format: The shape of the output: table, bullet list, JSON, word limit, tone.
  5. Constraints and criteria: What to avoid and how you will judge success.

This skeleton also matches the split between instructions, context and output format that Google describes in its Gemini prompting strategies. Therefore you can treat it as a universal frame, whichever model you use.

Does assigning a role really change the output?

Yes, although it is not a magic wand. A role influences vocabulary, level of detail and point of view. When you write "explain this as an experienced tax advisor", the model sounds more cautious and technical. When you write "explain it as you would to a high school student", it simplifies.

That said, a role alone never replaces context. If you write "You are an SEO expert" and then say nothing about the site, the industry or the goal, the model still produces generic advice. So I always back the role with concrete facts.

On the API side you usually place the role in the system message; Anthropic's documentation recommends exactly that. In chat interfaces, putting it at the top of your first message works fine.

How much context should you give?

The rule is simple: write down everything a smart new colleague would need to know. The model does not know your company, your customer or what failed last month. When you leave that out, it fills the gap with average internet text.

For example, when I ask for an ecommerce category description, I add the audience, price segment, what sets us apart from competitors, the brand voice and phrases to avoid. Those five lines move the result from "any shop" copy to copy that clearly belongs to one brand.

Still, do not inflate the context. Long documents unrelated to the task can distract the model. I filter context with one question: "Would the result change without this line?" If the answer is no, I delete it.

Also watch how fresh your context is. If last year's price list or an old campaign name stays in a template, the model treats it as today's truth. For that reason I review the fixed context blocks in our templates every quarter, so nobody multiplies outdated facts by accident.

How do you write clear and direct instructions?

The biggest jump in quality comes from swapping vague verbs for concrete tasks. Instead of "improve this text", write "cut every sentence below 20 words, switch to active voice, shorten the first paragraph to 50 words". Now the model no longer has to guess what "improve" means.

  • Say what to do rather than what not to do: instead of "no jargon", write "use words a 12 year old would understand".
  • Give numbers: "at most 80 words" instead of "short".
  • Define the order: if steps matter, use a numbered list.
  • Add the reason: "This text will appear on mobile, so paragraphs need to stay short."

Most people skip that last point. Anthropic's documentation notes that explaining the reason behind an instruction helps the model generalize better. Put simply, the model understands the goal instead of memorizing a rule.

In practice I test this by handing the prompt to a colleague who knows nothing about the topic. If they can do the job without asking questions, the prompt is clear enough. If they ask questions, I add the answers to the prompt. This small habit removes most misunderstandings before the model ever sees the request.

How do you control the output format?

Format is the part of prompt engineering with the fastest payoff. If you know how you will use the output, describe it in the prompt. For instance, if you will paste it into a spreadsheet, name the columns; if software will read it, write the JSON schema.

Moreover, a short template showing the desired shape works wonders. A pattern like "Headline: ... / Description: ... / Target keyword: ..." beats long explanations. OpenAI and Google also offer structured output options with JSON schemas on the API side.

Tone and style belong to format too. The style of your prompt tends to leak into the answer. So when I need a formal report, I write the prompt in a tidy, formal way; when I want flowing prose without bullets, I write the prompt as plain paragraphs.

How do XML tags and delimiters organize a prompt?

In long prompts, instructions, examples and data blur together. The model may confuse which line is an instruction and which is text to process. The fix is to separate sections with clear delimiters. Anthropic recommends XML tags for this, and OpenAI's guide also mentions Markdown headings and XML style delimiters.

In practice I use a structure like this:

  • <context> brand and audience information
  • <document> the raw text to process
  • <instructions> numbered steps
  • <output_format> the expected template

As a result, both the model and any colleague who edits the prompt later can see at once where everything sits. When the document changes, you only update the content of one tag. There is no official standard for tag names; as long as you stay consistent, you can choose your own.

What is the difference between a system prompt and a user prompt?

A system prompt is the persistent frame that applies to the whole conversation: role, tone, general rules, restrictions. A user prompt is the concrete request of the moment. Through an API you set this split with message roles; in the apps, ChatGPT custom instructions, Claude projects and Gemini Gems serve a similar purpose.

AspectSystem promptUser prompt
ScopeWhole conversation or applicationA single request
Typical contentRole, tone, brand rules, restrictionsTask, data, format for that job
How often it changesRarely, with version controlEvery message
Who writes itProduct owner or team leadThe end user
Impact of a mistakeSpreads to every outputAffects one output

The last row matters most. A small error in a system prompt spreads to hundreds of outputs. That is why my team keeps system prompts under version numbers and tests every change against a fixed set of inputs.

Where do few-shot and chain of thought fit into prompt engineering?

These two techniques are the best known tools in the prompt engineering kit, and I cover each in depth in separate articles. Here I only want to show where they sit.

Few-shot prompting means adding a few input and expected output examples to the prompt. The model infers the pattern from them. It helps most when you need consistent tone and format; however, examples that look too similar can lead the model to copy them.

Chain of thought means asking the model to reason step by step before answering. When the model writes out intermediate steps, you can also see where it went wrong. It can raise accuracy for calculations, comparisons and multi step logic. On the other hand, today's reasoning models already do this internally, so you may not need to request it every time.

What is prompt chaining and when should you use it?

Prompt chaining means splitting a big job into smaller prompts that feed each other. The output of each step becomes the input of the next. In one giant prompt the model may skip steps; in a chain, you check each step separately.

For a blog article, for example, we use this chain:

  1. Extract the search intent and the questions users ask.
  2. Build a heading outline from those questions.
  3. Write each section on its own.
  4. Review the full draft against the brand rules and list the fixes.
  5. Apply the fixes.

This setup has two big advantages. First, when something breaks, you rerun only that step. Second, the review step in the fourth stage lets the model catch problems in its own draft. Anthropic's documentation also recommends chaining for complex tasks.

How do temperature and other parameters affect results?

If you work through an API, some settings next to the prompt also shape the result. The best known is temperature. A low value gives more consistent and predictable output; a higher value gives more variety. So I prefer low values for data extraction or classification and higher values for brainstorming.

  • Temperature: the balance between variety and consistency.
  • Maximum output length: a limit high enough that the answer does not stop halfway.
  • Stop sequences: make the output end at a defined point.
  • Reasoning settings: on some models, a thinking budget or effort level.

Keep one thing in mind: not every model supports every parameter, and defaults change. For example, Google recommends keeping temperature at its default value for Gemini 3 models. Therefore I suggest checking the current documentation of your model before you touch a parameter.

How do you test and improve prompts?

Prompt engineering is not a write once job; the engineering part lives in the testing loop. I start with success criteria: "Is the output on brand, are the numbers correct, did the format hold?" Then I prepare ten to twenty realistic inputs for those criteria.

Next, I run two versions of the prompt on the same inputs and put the outputs side by side. One beautiful result can mislead you; what counts is how many of the twenty inputs meet the criteria. This way I get a concrete comparison like "eighteen out of twenty" instead of "I think it's better".

When you pick test inputs, do not rely only on easy cases. Add inputs with missing information, contradictions or unusual length, because a prompt usually shows its weak spots at the edges. For example, decide in advance whether the model should say "not enough information" when a product description comes in empty, and write that into the prompt.

Change one thing at a time as well. If you change the role, the format and the examples together, you will never know which change helped. Both OpenAI and Anthropic describe evaluation sets (evals) as a basic step for production grade applications.

What are the most common prompt mistakes?

In the teams I consult, I see the same mistakes again and again. The list below explains most of the lost time:

  • One line requests: "Write an Instagram post" with no context at all.
  • Conflicting instructions: "Keep it short" and "cover every detail" in the same prompt.
  • No stated goal: not saying who the text is for and what it should achieve.
  • No verification: publishing numbers, dates and sources without checking them.
  • Endless threads: expecting the model to remember the first rules after thirty messages.

The last two deserve special attention. Models can produce confident but wrong information, which people call hallucination. For that reason we verify every output that contains statistics, prices or regulations against a primary source. In long threads, the best fix is to open a new chat with a clear summary.

How should you structure prompts for long documents?

Current models can read very long texts in one go. You can hand over a contract, a report or hundreds of customer reviews. However, in long contexts the order of your prompt directly affects the result. Anthropic's documentation recommends placing long documents at the top and your question at the end.

I add two habits to that. First, I wrap each document in its own tag with a source name. Second, I ask the model to pull out the relevant quotes before answering. That way the model grounds its answer in actual sentences from the document, and I can see which part it used.

For instance, when I analyze a hundred reviews, I first say "quote every review that contains a complaint", then "extract five main themes from these quotes". This two step approach improves both accuracy and auditability. Moreover, when a wrong theme appears, I can trace it back to its quote right away.

How do you ask a model to review its own output?

A simple but strong technique is to have the model reread finished work against a checklist. The first prompt gives you the draft; the second says "review this text against these five rules and list every violation with its line number". In a third step, the model applies the fixes.

This works especially well when brand rules are strict. Banned phrases, character limits or mandatory legal notes often slip through on the first pass. In a separate review round, however, the model catches most of them.

Still, self review does not replace human review. A model can keep believing a number it invented. So I always add the line "flag every claim you cannot verify" to the checklist. The final read always stays with someone on my team.

How do you protect data privacy when writing prompts?

Everything you type into a prompt falls under the data policy of the service you use. Consequently, think twice before you paste personal data such as customer names, phone numbers, ID numbers or health details. Under laws like the GDPR, that responsibility stays with you.

  • Anonymize personal data with placeholders like "Customer A" or "City X".
  • Read the data usage terms of business plans; settings can differ from personal accounts.
  • Process confidential contracts and financial statements only in approved tools.
  • Write a clear team rule about which data may go into which tool.

I suggest putting these rules at the very top of your prompt library, so everyone using a template sees the same limits. I explain the broader frame for protecting data on your site in the website data security guide.

What do prompt engineering examples look like in marketing?

To make the theory concrete, here is a before/after example from a real type of workflow. Say we need headlines for a Google Ads campaign.

Before: "Write Google ad headlines for a dental clinic."

After: "You are a copywriter experienced in Google Ads search campaigns. Client: a dental clinic in Istanbul focusing on implants. Audience: patients over 40 who compare prices. Task: write 10 responsive search ad headlines of at most 30 characters. Three should focus on location, three on trust, four on easy booking. Avoid promises that break healthcare ad rules, price guarantees and the word 'best'. Return a two column table: headline and character count."

The second prompt is longer, yet it cuts the correction rounds almost to zero. Also, the character count column makes the result easy to check, although we still recount before launch. On the campaign side we test and filter these outputs as part of Google Ads management.

What does a before/after prompt look like for customer communication?

My second example comes from customer service, because a wrong tone here hits reputation directly. Imagine we need a reply to a complaint about a late delivery.

Before: "Write a polite reply to this complaint."

After: "You are the customer experience lead of an ecommerce brand. Reply to the complaint below. The customer received the order five days late and feels let down. Apologize first, own the delay without blaming the courier, and offer a concrete remedy: free shipping on the next order. Use at most 120 words and a warm but professional tone. Do not promise exact delivery dates or anything legally binding. End with a direct contact channel."

The second prompt carries the brand's policy and limits into the model. As a result, the reply is both more human and safer. Once you build this template, you can reuse it again and again by swapping only the complaint text.

How should you build prompts for SEO content?

In SEO content, the goal of prompt engineering is to tie the model to search intent and real expertise. Otherwise you get a copy of the average text on the web, which helps neither users nor Google. Google's E-E-A-T criteria evaluate exactly those experience and trust signals.

That is why my prompts always include the target keyword and intent, the questions users want answered, our own field observations and generic phrases to avoid. Feeding the keyword side with the keyword suggestion tool and checking the draft with the readability checker is a good start.

The model writes a draft; you add the experience. For SEO friendly content, I see AI as an assistant, not a replacement for the writer. If you also want visibility in AI search, add a GEO approach to your strategy.

How do you create a reusable prompt library for your team?

When a team starts using AI seriously, everyone writes their own prompts and quality starts to swing. The fix is a shared prompt library. We keep one template per task type: ad headline, product description, customer reply, meeting summary.

Each template has four parts: purpose, variable fields in square brackets, a sample output and the last test date. As a result, a new colleague gets the same quality on day one. And when a model update arrives, we know exactly which templates to retest.

You can also connect this library to your website or internal tools. In the article on using AI on your website I cover chatbot and content flows; the system prompts there need the same discipline.

Do you need different prompts for different models?

The core principles stay the same, but the fine tuning differs. ChatGPT, Claude and Gemini each stress slightly different things in their documentation. For example, Anthropic highlights XML tags, while Google discusses output prefixes and placing the question at the end of the prompt.

Reasoning models and fast models also behave differently. With reasoning models, a clear goal and clear constraints often work better than detailed "think through these steps in order" instructions. The OpenAI guide likewise treats these two model families separately.

My practical advice: when you move a prompt to another model, rerun your test set. A prompt that shines on one model can feel bloated or incomplete on another. Note the model name and version in your template too.

Vendors update their models often and retire older versions over time. Consequently, a note like "tested on this model on this date" helps you understand months later why quality dropped. When a new version ships, try it on a small test set first, then roll it out to the whole team.

Is prompt engineering a job or a skill everyone should learn?

I think it is both. Companies that embed AI in their products need dedicated expertise for system prompts, evaluation sets and model selection. Those roles usually work closely with software and product teams. Python helps a lot here; the Python AI roadmap explains where that path starts.

On the other hand, for everyday users prompt engineering is turning into a basic work skill, much like writing a good email. Marketer, salesperson or accountant, it makes no difference: whoever structures the request well does better work in less time.

In short, the title "prompt engineer" may change, but the skill will last. At its core it means thinking clearly and explaining clearly, and that keeps its value even when the tools change.

Where should you start with prompt engineering?

The best start is a task you do every day: a weekly report, a customer email or a social media caption. Write it with the five part structure, try it on five real cases and note the results.

Then improve the prompt one change at a time and save the best version as a template. Read the official guides as well; Anthropic, OpenAI and Google publish them for free and update them regularly. If you are curious about the coding side of AI, the article on AI coding tools for developers is a good next read.

If you want to place AI strategically inside your marketing, my team and I can build your content workflows together as part of SEO consulting. After all, a good prompt is a good brief, and a good brief is half of good work.

Frequently Asked Questions

Do I need to code to learn prompt engineering?
No, you do not need to code for prompt engineering in chat interfaces. Clear writing, giving context and testing the result are enough. However, if you plan to work through an API, embed system prompts in a product or build evaluation sets, basic Python knowledge helps a lot and lets you automate the whole process.
How long should a good prompt be?
A good prompt should be as long as it takes to describe the task fully; there is no fixed word count. A simple translation needs one sentence, while brand copy needs role, context, task, format and constraints. The measure is not length but how many gaps the model has to guess. Remove anything unrelated to the task.
Why does the same prompt give different results each time?
Language models work probabilistically, so the same prompt can produce different outputs. Parameters like temperature affect that variety. For consistency, define the format clearly, add a sample output and, if you use an API, apply the settings your model's documentation recommends. Test the prompt on several inputs to see the spread.
Should I write separate prompts for ChatGPT, Claude and Gemini?
The core structure can stay the same, but fine tuning may be necessary. The official guides from all three companies agree on clear instructions, context and format. Differences show up in delimiter preferences, parameter advice and how reasoning models behave. When you move a prompt to another model, retest it with the same inputs.
Does prompt engineering prevent hallucinations completely?
No, prompt engineering reduces hallucinations but does not remove them. Providing source documents, asking the model to rely only on them and allowing it to say it does not know all lower the risk. Still, you need to verify every output containing numbers, dates, prices or regulations against a primary source before publishing.
Does a small business need a prompt library?
Yes, even a small business benefits from one. Templates for the five to ten most repeated tasks make output quality independent of who writes the prompt. A simple shared document is enough: purpose, variable fields, sample output and last test date. Retest the templates whenever the model changes to keep quality steady.
#prompt engineering#prompt writing#artificial intelligence#ChatGPT#Claude#Gemini#system prompt#prompt chaining
Share:
Talha Aslan
Talha Aslan

Google Partner digital marketing expert. Hands-on with SEO, Google Ads, web design and e-commerce projects since 2012; every post here comes from that experience.

Next project

Let's talk about your project.

Your brief goes straight to Talha Aslan and team: strategy led by Talha, delivery by an experienced team. The first consultation is free; we listen and come back with a clear roadmap.

WhatsApp Call Now