What Is Chain of Thought Prompting? A Practical Guide With Examples

When you ask an AI model a complex question, the answer sometimes looks right at first glance, yet the math slips somewhere in the middle. Chain of thought prompting is the technique that addresses exactly this problem. In this guide I explain how it works, the research behind it, how it relates to today's reasoning models and how my team and I use it in marketing work.
What is chain of thought prompting?
Chain of thought prompting is a technique that asks a language model to write out intermediate reasoning steps before giving its final answer. The model breaks the problem into smaller steps, states each one explicitly and reaches the answer at the end of that chain, which usually improves accuracy on multi-step questions.
In practice, the idea mirrors how people solve problems on paper. For example, a math teacher wants to see the working, not just the result. Chain of thought asks the model to show its working in the same way.
There is an important distinction here. Chain of thought is not a type of model; it is a way of writing prompts. In other words, you ask the same model the same question in two different ways: once for a direct answer, once for step-by-step reasoning. The difference becomes clear on tasks that involve arithmetic, logic or planning.
Where did the chain of thought idea come from?
The paper that introduced the concept under its current name is "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" by Jason Wei and colleagues at Google, published in 2022. It showed that large models improve markedly on reasoning tasks when a few examples demonstrate the intermediate steps.
For their experiments, the researchers gave the model eight worked examples. Each one contained a question, then a step-by-step solution, then the answer. On a new question, the model imitated this pattern and produced its own intermediate steps.
The most striking finding was this: with the method, the 540-billion-parameter PaLM model beat the best fine-tuned GPT-3 result of the time on GSM8K, a benchmark of grade-school math word problems. The authors also noted that the benefit emerged with model scale. Smaller models often wrote fluent but wrong steps, so at that point the technique worked mainly on large models.
How does a reasoning chain work inside the model?
First, remember that a language model generates text token by token. Each new token depends on all the text before it. So when the model writes the answer immediately, it has nowhere to "hold" intermediate calculations. When it writes the steps out, however, each step becomes context for the next one.
Put simply, here is an analogy. Multiplying 37 by 24 in your head is hard; if you write 37 times 20 first and then 37 times 4, the task gets easier. Chain of thought hands the model that sheet of paper. Every written intermediate result becomes the basis for the next prediction.
On the other hand, this mechanism offers no guarantee. If the model makes a mistake in one step, later steps build on that mistake. Still, making the error visible is a big advantage, because you can read the chain and find where the problem started.
How does a standard prompt differ from a CoT prompt?
An example shows the difference best. From here on I will also use the common abbreviation CoT. The table below summarises the kinds of responses I get when I ask the same question in both ways.
| Feature | Standard prompt | CoT prompt |
|---|---|---|
| Request | "Give me the result." | "Write the steps first, then give the result." |
| Output length | Short | Longer, with intermediate steps |
| Accuracy on multi-step questions | Can be lower | Usually higher |
| Debugging | Hard, the reason is hidden | Easy, the faulty step is visible |
| Cost and latency | Low | More tokens, more time |
| Suitable tasks | Simple facts, classification | Calculation, logic, planning, comparison |
The takeaway is clear: CoT adds value on questions that require thinking, not on every question. Therefore, rather than blindly adding "think step by step" to every prompt, I recommend that you look at the type of task first.
What is few-shot CoT?
The original method Wei and colleagues used is known as few-shot CoT. You give the model several worked examples and write the reasoning steps explicitly in each one. The model then imitates that way of thinking on the new question.
Few-shot prompting itself is a separate topic; here I focus only on the part that adds reasoning. The key difference is simple: ordinary few-shot examples contain only input and output, while CoT examples place a rationale between the two.
In practice, I set up the example template like this:
- Question: a real business question, for example the cost per customer of a campaign.
- Reasoning: which data you use, which operation you apply and the intermediate results.
- Answer: a single, clear line.
Above all, the quality of the examples matters here. Moreover, the more consistent the rationale format in your examples, the more orderly the chain the model produces on a new question.
What is zero-shot CoT and "Let's think step by step"?
Few-shot CoT works, but preparing examples takes effort. In 2022, Takeshi Kojima and colleagues showed a simpler route in "Large Language Models are Zero-Shot Reasoners": add the sentence "Let's think step by step" to the end of the question, with no examples at all.
Still, that single sentence made a surprising difference in the paper's experiments. With an InstructGPT model, accuracy on MultiArith rose from 17.7% to 78.7%, and on GSM8K from 10.4% to 40.7%. In other words, the authors unlocked the model's reasoning with one trigger phrase and no hand-written examples.
The paper also proposed a two-stage setup. In the first stage, the model writes its rationale after the trigger phrase. In the second stage, you append a phrase such as "Therefore, the answer is" and the model extracts the result. I use a similar pattern in daily work: "Think step by step, then write the final answer on one line."
In short, zero-shot CoT is the fastest form of the technique to try. However, on complex, domain-specific tasks, well-chosen examples still deliver more consistent results.
What is self-consistency and how does it combine with CoT?
Self-consistency is a complementary method proposed by Xuezhi Wang and colleagues in a 2022 paper. Instead of a single chain, you ask the model to produce several different chains for the same question and then pick the most frequent answer.
The logic is intuitive, because a hard problem usually has more than one correct solution path, while wrong paths scatter across different wrong answers. As a result, a majority vote tends to surface the correct answer. In the paper, the method improved the CoT result on GSM8K by 17.9 percentage points.
In practice, self-consistency works in these steps:
- You run the same prompt several times with a temperature above zero.
- Next, extract the final answer from each run.
- Finally, pick the most frequent answer; if the answers scatter widely, you revisit the question.
The price, however, is cost. Producing five chains means spending roughly five times the tokens. That is why I use self-consistency only for decisions where a mistake is expensive.
How does chain of thought prompting relate to reasoning models?
Today OpenAI, Anthropic and Google offer reasoning models that think internally before they answer. These models moved the chain of thought idea from the prompt level to the model level. So in most cases you no longer need a separate sentence to make the model think.
According to the official documentation, the picture looks like this. OpenAI's reasoning guide tells you to control the amount of thinking with the reasoning.effort parameter; the reasoning tokens stay invisible in the API but count as billed output tokens, and you can request a summary. Google's Gemini thinking documentation offers a similar control through the thinking_level parameter and states that thinking tokens are part of the price. Anthropic, for its part, highlights adaptive thinking and the effort setting for current Claude models.
So the reasoning chain has not disappeared; it has simply moved. You used to trigger the chain yourself; now the model builds its own chain, and you manage its depth with a parameter.
Do you still need to say "think step by step" to reasoning models?
In most cases no, and sometimes it even backfires. Anthropic's prompting best practices page says that when thinking is on, general instructions such as "think thoroughly" often produce better reasoning than a hand-written step-by-step plan. The same page recommends classic CoT prompting as a fallback when thinking is off.
OpenAI's documentation follows a similar line: give the model the task, the constraints and the desired output format clearly, and treat the effort setting as a tuning knob rather than the main way to recover quality.
My practical rule is this: with a reasoning model, I do not dictate the thinking process. Instead, I describe the goal, the data, the acceptance criteria and the output format. With a fast, cheap model that has no thinking feature, I still use the classic CoT pattern.
How does CoT differ from prompt chaining?
Many people confuse the two because of their names. In CoT there is a single prompt, and the model writes its reasoning inside the same answer. In prompt chaining, you split the work into several separate requests; the output of one step becomes the input of the next.
For example, when preparing a blog post you can run topic research in the first request, a draft in the second and an edit in the third. That is a prompt chain. Inside each step you can still ask the model to reason step by step, so the two techniques do not exclude each other.
Anthropic's documentation makes this point too: current models handle most multi-step reasoning internally, yet splitting prompts into separate calls remains useful when you need to inspect intermediate outputs or enforce a specific pipeline. I prefer chaining for work that needs an audit trail.
Are there techniques that grew out of CoT?
Yes, researchers built many methods on the reasoning chain idea after 2022. You do not need to master all of them, but knowing what the names mean makes your work easier.
- Least-to-most prompting: splits a problem into sub-problems first, then solves them in order from the easiest.
- Tree of thoughts: explores several branches of reasoning instead of a single line and prunes weak branches.
- ReAct: alternates reasoning steps with tool calls such as search or calculation.
- Plan-and-solve: asks the model to write a plan first, then to execute it step by step.
In practice, you meet most of these variants inside an agent framework or in the built-in behaviour of reasoning models. That is why I recommend the basic technique to marketing teams first and the variants only when a real need appears.
How do you measure whether a reasoning chain actually helps?
Instead of deciding by feel, I suggest building a small test set. Pick ten to twenty real business questions whose correct answers you know, for example cost calculations from past campaigns. Then run each question with both prompts.
When comparing, I look at three things: the share of correct answers, response time and token usage. If accuracy rises clearly and the cost stays acceptable, the technique becomes permanent for that task. If there is no gain, solving the task with a simpler prompt makes more sense.
Also, rerun the test set whenever you switch models. A prompt that works on one model may behave differently, especially after a move to a reasoning model. This small discipline bases your AI decisions on measurement rather than guesswork.
What role does a reasoning chain play in AI agents?
Put simply, agents are AI systems that take several steps and use tools to reach a goal. For example, an agent reads an ads report, runs a calculation and then drafts a summary email. At every step, reasoning about what to do and why is critical for choosing the right tool.
That is why the reasoning chain works like a hidden backbone in agent design. Current reasoning models run this reasoning internally; your job is to define the goal, the permissions and the stopping condition clearly.
My advice is to log a summary of the rationale for an agent's important decisions. When something goes wrong, you can then trace back which step was faulty.
Which tasks benefit from step-by-step reasoning?
The technique creates value on any task that requires more than one intermediate step. In the research, the strongest effect showed up on arithmetic word problems, commonsense logic questions and symbolic manipulation.
In practice, I group the business tasks I encounter like this:
- Budget and cost calculations: cost per click, customer acquisition cost, profit margin.
- Comparative decisions: choosing between two offers, two campaigns or two price options.
- Rule checks: whether a text follows the brand guide or an advertising policy.
- Planning: content calendars, launch steps, task lists with dependencies.
- Data interpretation: ruling out possible causes of a drop in a report one by one.
In short, what these tasks share is simple: the path to the right answer matters as much as the answer itself. Seeing that path exposes both the model's errors and gaps in your own data.
When does chain of thought prompting not help?
Every technique has limits, and knowing them for CoT saves time and money. On single-step factual questions, such as asking for the capital of a country, writing out steps only makes the answer longer.
In addition, some studies show that step-by-step reasoning can hurt performance on certain tasks that rely on quick, intuitive judgement. Humans experience something similar: trying to explain step by step how you recognise a face makes the task harder.
The third limit is missing knowledge. If the model lacks the necessary information, a long chain only produces a more convincing wrong answer. In other words, CoT does not fill data gaps; it organises reasoning.
The fourth limit is creative work. Forcing a strict logic chain while generating slogans or campaign ideas often yields safer but blander ideas. For such work, free idea generation followed by step-by-step filtering gives better results.
Finally, there is latency and cost. In scenarios that need fast replies, such as a live chatbot, long chains hurt the user experience. A better design is to generate the rationale in the background and show the user only the result.
Does the written chain show how the model really thinks?
Not always. In fact, researchers call this the faithfulness problem. A model can write a plausible rationale even when it actually reached its answer through a different cue.
In practice this means the chain is a powerful audit tool, but it does not count as proof by itself. For example, a model may explain with reasons why an ad text complies with a policy; you still need to do the final check against the official policy page.
Therefore, I read the chain like an auditor. At each step I check where the numbers come from, whether the assumptions are reasonable and whether the reasoning skips anything. That way I use the model as an assistant that shows its work, not as an oracle.
How do you write a good CoT prompt?
A good CoT prompt has four parts: context, task, reasoning instruction and output format. When I follow the order below, results become noticeably more consistent:
- Give context: business type, target audience, the numbers you have.
- Define the task in one sentence: "Which campaign should get more budget?"
- Add a reasoning instruction: "First list the data, then calculate each option, then compare."
- Separate the output: request the rationale in one section and the final answer in another.
- Add a check step: "Verify your calculations once more before answering."
Above all, separating rationale and answer matters. Anthropic recommends structured tags for this separation when thinking is off. As a result, you can extract only the answer without pasting the rationale into your report.
Language consistency matters as well. If you write part of the prompt in one language and part in another, the model may reason in one and answer in the other. So state the output language at the start of the prompt, and use the same terms in every prompt.
What are examples of CoT in marketing?
To make the theory concrete, here are three scenarios my team and I use often. The numbers are illustrative and do not belong to a real client.
First, consider an ad budget. The prompt reads: "We have a monthly budget of 3,000 USD. Campaign A costs 0.60 USD per click with a 3% conversion rate; campaign B costs 0.90 USD with 5%. First calculate the cost per conversion for each, then justify how we should split the budget." The model finds 20 USD for A and 18 USD for B, concludes that B is more efficient and adds that you also need to check B's volume ceiling.
Second, consider content strategy. I ask the model to group a keyword list by search intent first, then justify which page type each group fits, and finally rank the priorities. The step-by-step approach turns a messy list into a readable plan.
The third scenario is headline selection, which also benefits from explicit reasoning. I ask the model to rate three headline candidates separately for audience fit, clarity and match with the search phrase. Then I compare the result with the headline analyzer.
How can you use step-by-step reasoning in business processes?
Outside marketing, the technique pays off as well. For example, in customer service, having the model check step by step whether a return request fits the policy reduces arbitrary decisions. The model checks the purchase date first, then the product condition and finally any exceptions.
Likewise, sales teams can use it to prioritise incoming leads. The model evaluates budget, urgency and fit signals one by one and gives a score with reasons. That way the sales rep sees where the score comes from and can challenge it.
In finance and operations, offer comparison, stock planning and pricing scenarios stand out. If you want to build this kind of work into your website, the article on using AI on your website is a practical starting point. I cover the automation side in AI in web design and automation.
What are the most common CoT mistakes?
Meanwhile, when I train teams, I see the same mistakes again and again. Knowing them is the shortest path to real value from the technique:
- Adding "think step by step" to every prompt and raising costs for nothing on simple tasks.
- Not separating rationale from answer, then hunting for the answer inside the text.
- Treating the chain as proof and not verifying where the numbers come from.
- Forcing a rigid step list on a reasoning model and limiting its own thinking.
- Asking for a long chain with incomplete data and getting a convincing but wrong result.
Most of these mistakes disappear with one habit: before writing the prompt, ask whether the task truly has several steps. If yes, request the chain; if not, go straight for the answer.
How does step-by-step reasoning affect SEO and content work?
On the content side, CoT adds value in planning more than in the writing itself. For example, when I ask the model to analyse reader questions, search intent and gaps in competing content in order before drafting an outline, the result is a more coherent structure.
On the other hand, publishing the model's rationale directly is a mistake. Readers want a clear, verified conclusion, not raw reasoning. So I keep the chain as an internal working note and edit the published text separately. I explain how I structure content for AI search visibility in how to write content for AI Overviews, and the bigger picture in is SEO dead.
After that, a human does the final check. I measure flow with the readability checker and trace every claim back to its source.
How do you protect data privacy when using reasoning chains?
However, step-by-step reasoning encourages you to give the model more data. Teams naturally add more figures, customer details and internal reports to the prompt. At this point you need to review your privacy rules.
- Anonymise personal data such as names, phone numbers and emails before adding them to a prompt.
- For business use, read the provider's data retention and training policies on its official pages.
- If you log rationale output, decide who has access to those logs.
That way you gain the benefit of the technique without taking unnecessary risks. Remember, the longer the chain, the more data flows through it.
Where should you start?
After all this detail, let me reduce the work to a simple sequence. First, pick a multi-step task from your weekly work, such as splitting a campaign budget. Then ask the same question directly and with a CoT prompt, and compare the two answers.
If the difference is clear, turn that prompt into a template and share it with your team. If you use a reasoning model, sharpen the goal and acceptance criteria instead of the thinking instruction, and manage depth with the effort setting. For critical decisions, generate several chains with self-consistency and look at the majority.
To check ad budget math quickly, try the Google Ads budget calculator. If you want to bring AI into your ads and search strategy, my team and I apply these methods in Google Ads management, SEO consulting and social media management. To discuss your own scenario, write to me through the contact page.




