Artificial Intelligence

ChatGPT Prompts for Technical SEO: Templates for Robots.txt, Redirects, Schema and Speed

Talha Aslan 19 min read

ChatGPT prompts for technical SEO turn the model into a fast assistant for recurring technical work, from robots.txt reviews to redirect maps. However, the model does not crawl your site; it interprets the files and exports you give it. In this guide I share the prompt templates I use on client projects, ready to copy and adapt, along with how I validate each result.

What are ChatGPT prompts for technical SEO?

ChatGPT prompts for technical SEO are structured instructions that ask a language model to analyse crawling, indexing, redirect, structured data and speed issues. A good prompt defines the role, the input, the rules and the output format. You always validate the result with official tools before changing anything on a live site.

The critical part, however, is the input. The model does not connect to your site and browse pages; even with web search on, it does not run a full crawl. So you give it your robots.txt file, a Lighthouse report, a crawler CSV export or an excerpt of your server logs. Then the model reads that data, finds patterns and suggests fixes.

In practice, I covered general technical SEO concepts in what is technical SEO. This article focuses on practice: which prompt you write for which task, and when you can trust the output.

Adapt the details in each template to your site. For example, when you add your industry, your CMS and your server type at the start of the prompt, the model gives more accurate suggestions. Whether you run WordPress, Shopify or custom software directly affects whether a proposed fix is realistic.

What can ChatGPT do in technical SEO, and what can't it do?

Also, the model is strong at text and pattern work. For example, it groups long URL lists, writes regular expressions, drafts JSON-LD, explains .htaccess rules and translates error messages into plain language.

On the other hand, it is weak at anything that needs live data. It cannot know whether a page is actually indexed, how Googlebot rendered it or what real users experience in terms of speed. That information lives in Search Console, PageSpeed Insights and your server logs.

In short, the right division of labour looks like this: official tools and crawlers collect the data, the model organises and interprets it, and you make the final decision. In practice, the table below summarises this split for every prompt in this guide.

TaskChatGPT's roleValidation tool
robots.txt reviewReads rules, finds conflictsSearch Console robots.txt report
Redirect mapMatches old and new URLs, writes rulesRedirect checker, browser
Schema generationDrafts JSON-LDRich Results Test
hreflang checkFinds missing return linksCrawler, Search Console
Speed diagnosisPrioritises Lighthouse outputPageSpeed Insights field data
Log analysisSummarises bot patternsReverse DNS verification

How do you build a good technical SEO prompt?

I use the same four parts in every template, because this structure narrows the space where the model can guess.

  1. Role: a frame such as "You are an experienced technical SEO specialist".
  2. Input: the file or data itself, clearly delimited.
  3. Rules: "rely only on the data provided", "write CHECK if unsure", "do not suggest anything that contradicts Google's official documentation".
  4. Output format: a table, a bullet list or a code block you can copy directly.

Also ask the model to add a risk level next to each suggestion. That way warnings like "this rule could block the whole site from crawling" do not slip past you. General prompt writing techniques are a separate topic; here I focus only on patterns specific to technical SEO.

Watch the input size too. With very large files the model can miss details towards the end. Therefore I filter crawler output first and send only the columns I want to review and the rows with problems. For example, I send only rows returning 3xx and 4xx rather than the whole site.

Which prompt should you use for a robots.txt review?

A robots.txt file is small; however, one wrong line can block your whole site from crawling. Here is my template:

Prompt: "Review the robots.txt file below. For each User-agent group, list the blocked paths in a table. Flag rules that block CSS, JavaScript or image files. Show conflicting Allow and Disallow lines. Note if the Sitemap line is missing. Give a risk level for each finding."

When you read the output, remember one rule: robots.txt controls crawling, not indexing. Google's robots.txt documentation states clearly that a blocked page can still get indexed if other sites link to it. Also, the model sometimes suggests robots.txt as a way to remove pages from the index; reject that suggestion.

If you are creating a new file, build the basic skeleton with the robots.txt generator and ask the model to check only your custom rules.

How do you get robots.txt rules for AI crawlers?

This is, in fact, one of the most frequent technical questions lately. In practice, you can ask the model to write the rules, but always check bot names against official sources, because the model can produce outdated or wrong names.

According to official documentation, OpenAI uses GPTBot for training and OAI-SearchBot for ChatGPT's search features. Anthropic runs ClaudeBot, and Perplexity runs PerplexityBot. Google-Extended, meanwhile, is not a separate crawler; it is a token that controls whether your content is used for training and grounding Google's Gemini models. Blocking it does not affect your ranking in Google Search.

Prompt: "Write robots.txt groups that block the whole site for GPTBot and Google-Extended, but allow only the /blog/ folder for OAI-SearchBot. Keep all other rules in my current file. List the changed lines separately."

I discussed the strategic side of this decision in technical SEO after AI. To summarise your site for AI systems, you can also try the llms.txt generator.

What should ChatGPT prompts for technical SEO look like for redirect maps?

In site migrations, matching old URLs to new ones takes the most time. ChatGPT really speeds this up, because it can match two lists by title and path similarity.

Prompt: "Column A contains old URLs and page titles, column B contains new URLs and titles. For each old URL choose the best new URL. Write CHECK for matches you are unsure about. List pages without a counterpart separately. Return a three column table: old, new, confidence."

Then a second prompt writes the rules: "Convert this table into 301 rules for Apache .htaccess. Warn me about any rule that would create a chain." If you run Nginx, write the same prompt for Nginx syntax.

That said, validation is essential. Test the rules on staging first, then use the redirect checker to confirm sample URLs reach the right target in a single hop. For the full migration, see the website migration SEO checklist.

Which prompt should you use to write regular expressions?

Search Console filters, Google Analytics segments and redirect rules all need regular expressions. Most people, however, do not know the syntax by heart, while the model usually produces a correct draft.

Prompt: "Write a regular expression for the Search Console performance report. It should match queries that start with a question word: how, what, why, which, when. Keep in mind that Search Console uses RE2 syntax. Give the expression plus three example queries it matches and three it does not."

The last sentence matters, because it gives you test cases. Asking for positive and negative examples makes it easy for you to test the logic. RE2 does not support some advanced features such as lookbehind, so stating the syntax upfront prevents errors.

Before applying the expression to a report, test it on a few real queries. If you see an unexpected result, give that query to the model and ask it to fix the expression.

How should you write ChatGPT prompts for technical SEO schema markup?

Generating JSON-LD is one of the things the model does best. Still, there is a fabrication risk here too: the model can add ratings, prices or reviews that do not appear on the page. That violates Google's structured data guidelines.

Prompt: "Generate schema.org JSON-LD for the page content below. Use only information stated explicitly in the text. Leave out fields that are not in the text and list them separately. State Google's required and recommended properties for this type separately. Page type: local business service page."

Next, validate the output with Google's Rich Results Test. Even if the test shows no errors, check yourself that the markup matches the visible content exactly.

For a quick start, use the schema generator. Also, I explained the logic of structured data in detail in what is schema markup.

How do you find hreflang errors with ChatGPT?

On multilingual sites the most common problem I see is missing return links. Page A points to B, but B does not point back to A. In that case Google may ignore the annotation.

Prompt: "In the table below each row is a page and its hreflang tags. Check whether every pair is reciprocal. Flag invalid language or region codes. List groups missing x-default. Note any page without a self referencing hreflang line."

First, you export the input from a crawler as CSV. Then the model reads the table and finds the gaps. The source of truth is Google's localized versions documentation; if a suggestion looks questionable, check it there.

If you need to regenerate the tags, the hreflang generator speeds things up. For the basics, read what is the hreflang tag.

How do you run a canonical and duplicate content audit?

Export URL, title, canonical and status code columns from your crawler. In my experience, this table is one of the most efficient inputs you can give the model.

Prompt: "In this table, list pages whose canonical tag points to a different URL. Flag rows where the canonical target returns 404, 301 or noindex. Group URLs that share the same title. Summarise the canonical status of parameterised URLs separately."

In practice, the model handles cross checks like these quickly. However, a canonical is a hint, not a directive; Google may choose a different URL as canonical. So compare the model's findings with the Google selected canonical in Search Console's URL Inspection tool.

Moreover, groups of duplicate titles often point to a deeper structural problem. In that case you need to solve it at template level rather than page by page.

Which prompt works for a sitemap check?

The purpose of an XML sitemap is to show Google your important, indexable URLs. Yet on many sites the sitemap is full of redirected, noindexed or 404 addresses.

Prompt: "Below are the URLs in my sitemap along with the status code, indexability and canonical data my crawler returned for them. List URLs that should not be in the sitemap, with reasons. Separately, show indexable pages returning 200 that are missing from the sitemap."

As a result, this comparison reveals the gap between your sitemap and the real site structure. To produce a clean file, use the XML sitemap generator.

After the fix, resubmit the file in Search Console's Sitemaps report and check the number of discovered pages a few days later.

How does ChatGPT help with Core Web Vitals diagnosis?

Lighthouse reports, for instance, are long and technical. Also, the model is genuinely useful at reading them and setting priorities. However, Lighthouse is lab data; the signals Google uses for ranking rely on real user data.

Prompt: "In the Lighthouse JSON output below, find the audits related to LCP, CLS and INP. For each issue, write the likely root cause, the affected metric and the estimated implementation effort. Return a table sorted by highest impact and lowest effort."

INP is the interaction metric that replaced FID in March 2024; the web.dev INP documentation sets the good threshold at 200 milliseconds or less. If the model brings FID advice from older articles, remind it of this change.

Track the effect of fixes through field data in PageSpeed Insights. I covered how speed affects rankings in how does site speed affect SEO.

How do you get ChatGPT to analyse server logs?

Log analysis is the only source that shows what Googlebot actually crawls on your site. Because the files are large, give the model a summarised excerpt rather than the raw log.

Prompt: "Below are lines from the last seven days of my access log with a Googlebot user agent. Summarise the most crawled directories, URLs returning 4xx and 5xx, the share of requests spent on parameterised URLs and important sitemap pages that were never crawled."

There is a trap here, however: user agents are easy to fake. Google recommends verifying Googlebot with a reverse DNS lookup or by comparing against its published IP ranges. So verify the lines before relying on the summary, or at least note that the data is unverified.

IP addresses in log files can count as personal data. Mask or remove the IP column before sending anything to the model.

Which prompt should you write for JavaScript and rendering issues?

On JavaScript heavy sites the key question is whether the HTML Google sees matches the page users see. In practice, the model cannot test this itself, but it can compare two outputs.

Prompt: "Below are the raw HTML source and the rendered HTML of the same page. List differences between the two versions in terms of title, meta description, canonical, H1, internal links and main body text. Flag critical elements that appear only in the rendered version."

In practice, you can take the rendered HTML from the live test in Search Console's URL Inspection tool. If critical content arrives only through JavaScript, discuss server side rendering or prerendering options with your team.

This comparison prevents surprises, especially on sites moving to a new framework. For example, during a React or Vue migration you can catch internal links that only work through click handlers and carry no real anchor element. Google follows only links with an href attribute reliably.

Which prompt suits a meta tag and title audit?

On a site with hundreds of pages, finding missing, overly long or duplicate titles by hand is tedious. You can give the crawl table to the model and also ask for bulk rewrites.

Prompt: "This table contains URL, title and meta description. List titles longer than 60 characters, empty descriptions and exact duplicate titles. For problem rows, write new suggestions that keep the page topic. Do not append the brand name to titles."

Still, suggestions are only a starting point. Especially on product pages, add the main keyword of each page as a separate column to stop the model drifting into generic phrases.

Also ask the model to state the character count for each suggestion. Language models sometimes miscount characters, so check the number with a tool anyway. Use the meta tag generator and the Google SERP preview for length and appearance.

How do you prioritise broken links and internal linking issues?

Crawlers often list hundreds of broken links. You cannot fix them all at once, so you need priorities.

Prompt: "This table contains source page, broken target URL and anchor text. Group the broken targets. For each group, state how many pages link to it. Suggest a live replacement URL from the list of live URLs below; if none fits, write 'remove'."

Also, this prompt pushes the broken target with the most internal links to the top. As a result, a single redirect repairs dozens of links at once.

For regular checks, then, use the broken link checker. For a general health overview, the SEO checker offers a quick start.

How do you get the Search Console indexing report interpreted?

The Page indexing report in Search Console groups the reasons why pages are not indexed. Statuses such as "Crawled, currently not indexed" or "Alternate page with proper canonical tag" confuse many site owners.

You can export the sample URL list of this report and give it to the model. Prompt: "Below are URLs from the 'Crawled, currently not indexed' group in Search Console. Group them by path structure. For each group, list likely causes: thin content, duplicate template, parameters, pagination. Explain which group to review first and why."

The model does not make a definitive diagnosis here; instead, it splits hundreds of URLs into meaningful clusters. For instance, you quickly see that most problems come from tag pages or filter parameters. Then you inspect a few examples from each cluster in the URL Inspection tool.

Also read the meaning of status names from Google's help pages, not from the model. Report names change from time to time and the model may use old ones.

Which prompt helps with parameterised URLs and crawl budget?

On e-commerce sites, filter, sort and session parameters create thousands of URL variations. On large sites this can reduce the time Googlebot spends on important pages.

Prompt: "Below is a list of parameterised URLs from my crawler. List each parameter by name, count how many URLs contain it and guess its function: filter, sort, pagination, tracking or session. For each parameter suggest an approach: canonical, noindex, robots.txt block or leave as is. Write CHECK on rows you are unsure about."

This output therefore makes a good draft for a meeting with the development team. However, blocking parameters is hard to reverse. If you accidentally close a filter page that has search demand, you lose traffic.

Google notes that crawl budget mainly matters for very large or frequently updated sites. If your site has a few hundred pages, do not prioritise this analysis.

How should you organise your prompt library?

Rather than using these templates once and forgetting them, turn them into a library. In practice, I keep prompts in a simple table:

  • Prompt name and purpose
  • Required input and format (CSV, JSON, plain text)
  • Prompt text and version number
  • Validation tool and check step
  • Last update date and known weak spots

The version number, specifically, is a small but important detail. When you change a prompt, you can compare old results with new ones. Also, when Google updates a rule, you quickly find which prompts it affects; the switch from FID to INP, for example, meant updating every speed prompt at once.

Next, share the library with your colleagues. That way everyone works with the same rules and results do not vary from person to person.

Which mistakes should you avoid with ChatGPT prompts for technical SEO?

In the audits my team and I run, we see recurring mistakes in AI generated suggestions:

  • Suggesting robots.txt to remove pages from the index; the right route is a noindex tag on a crawlable page.
  • Adding ratings or reviews to schema that the page does not show.
  • Giving speed advice based on the retired FID metric.
  • Using bot names that do not exist or outdated product names.
  • Adding new rules without spotting redirect chains.
  • Trusting log data without verifying the Googlebot user agent.

Above all, these mistakes share one thing: the model produces an answer that sounds reasonable but contradicts current official rules. That is why I add "do not suggest anything that contradicts Google Search Central documentation; say so if unsure" to my prompts as standard.

Which process should you follow to validate the output?

After all, technical changes are expensive to reverse. A faulty robots.txt or a wrong redirect rule can cost weeks of traffic. Therefore I never push model output straight to production. Also, I apply these four steps to every change.

  1. Compare the suggestion with the official documentation.
  2. Apply it on a staging environment.
  3. Validate it with the relevant Google tool: Rich Results Test, URL Inspection, robots.txt report.
  4. After going live, monitor the related Search Console report for a few weeks.

Also log the prompts and the model's answers in a change log. When something breaks, you quickly find which change came from where.

Which data should you not share with ChatGPT?

Technical SEO data can be more sensitive than it looks. Server logs contain IP addresses, and configuration files hold internal server paths and sometimes passwords.

In practice, I follow one rule: I give the model only data that is already public or that I have anonymised. robots.txt and sitemaps are public anyway. In logs, I mask IPs; in .htaccess files, I strip secret values.

On business accounts, check the data usage settings. For sensitive projects, the API or enterprise plans offer extra assurances, such as not using data for training by default. In every case, though, the safest data is the data you never send.

Whose work do these prompts make easier?

The biggest gain goes to specialists who know technical SEO but spend too much time on repetitive tasks. In practice, the model does not decide for them; it organises data fast and produces the first draft.

Beginners benefit too, but on one condition: they use the model's explanations as a learning tool and check every suggestion against official documentation. Otherwise they risk applying wrong information to a live site.

For developers, prompts create a shared language. The SEO team defines the problem in a prompt, the model drafts the code, and the developer reviews and implements it. As a result, the back and forth between teams shrinks.

Consultants gain as well, because client reports come together faster. Still, every finding in a report needs verified data behind it. Pasting the model's sentences straight into a report damages trust, especially when they contain a wrong technical claim.

When should you hand this work to a professional?

In projects such as site migrations, domain changes, moves to a multilingual structure or serious crawl budget issues, the margin for error is small. Prompts speed up this kind of work, but an experienced person needs to own the process.

Within our SEO consulting work, my team and I use these templates together with our own validation steps. If you need to build the technical foundation from scratch, we plan SEO requirements at the very start of the web design process.

If you prefer to go on your own, start with a single task. For example, use only the robots.txt and sitemap prompts this week, validate the results and then move to other templates. At the end of the first week, note which prompt saved you time and which needed too many corrections.

Frequently Asked Questions

Can ChatGPT crawl my site and run a technical SEO audit?
No, it cannot run a full crawl. The model does not browse your site like a crawler and cannot know how Google renders your pages. The right approach is to give it your crawler's CSV export, robots.txt file or Lighthouse report. The model interprets that data, and you validate the result with official tools like Search Console.
Can I remove a page from Google with robots.txt?
No. robots.txt controls crawling, not indexing. A blocked page can still appear in the index if other pages link to it. To remove a page from the index, use a noindex tag and do not block the page in robots.txt, so Google can see the tag. For urgent cases, the Search Console removals tool offers a temporary fix.
Can I use schema code generated by ChatGPT directly?
You need to validate it first. The Rich Results Test checks syntax and required fields. Also make sure yourself that the markup matches the visible content exactly. The model sometimes adds ratings, prices or reviews that do not appear on the page; that breaks Google's structured data guidelines and must be removed.
Does blocking AI crawlers affect my rankings?
Blocking bots such as GPTBot or ClaudeBot does not affect your Google Search rankings. According to Google, blocking the Google-Extended token does not change Search ranking either. However, if you block search bots such as OAI-SearchBot, your visibility in the related AI search products may drop. Decide based on your goals and monitor the effect.
Is it safe to send log files to ChatGPT?
Do not send them raw. Logs contain IP addresses, which can count as personal data. Mask the IP column first, filter only the lines you need and send a summarised excerpt. On business accounts, check the data usage settings. Always strip sensitive configuration values before sharing anything with a model.
Which ChatGPT version should I use for technical SEO?
A current version with file upload and long context support makes the work easier, because it can read CSV and JSON files directly. Model names change often, so check OpenAI's current documentation. Whichever version you choose, the step of validating output with official Google tools stays the same.
  • ChatGPT
  • technical SEO
  • prompts
  • robots.txt
  • schema
  • Core Web Vitals
  • AI
Share:
Talha Aslan

Google Partner digital marketing expert. Hands-on with SEO, Google Ads, web design and e-commerce projects since 2012; every post here comes from that experience.

Next project

Let's talk about your project.

Your brief goes straight to Talha Aslan and team: strategy led by Talha, delivery by an experienced team. The first consultation is free; we listen and come back with a clear roadmap.