How to Use AI in Screaming Frog: Prompts, Embeddings and Semantic Similarity Guide

Screaming Frog AI features turned the crawler I have used for technical audits for years into a content analysis tool. In this guide I explain how I connect the SEO Spider to AI providers, which tasks actually save time and where you should stay careful. I checked every menu name against Screaming Frog's official user guide and release notes.
What is Screaming Frog AI and what does it do?
Screaming Frog AI is the SEO Spider's ability to connect to AI services such as OpenAI, Gemini, Anthropic and Ollama through their APIs and run your own prompt against every page it crawls. As a result, you can generate alt text, classify intent, detect language and analyse semantic similarity across thousands of URLs in one crawl.
A classic crawl gives you structural data: titles, meta descriptions, status codes and canonicals. The AI layer, on the other hand, interprets what a page actually says. For example, it can label whether a product page serves informational or transactional intent. In other words, the structural audit and the content audit finally meet in the same table.
The practical value of that merger is large. I used to export crawl data, run it through an AI model in a separate spreadsheet and then map the answers back to URLs. That process took hours and invited matching errors. Now the answers land directly on the relevant URL row, and you can filter and sort them together with every other crawl column.
Everything in this article sits inside a wider technical SEO framework. I recommend reading my piece on technical SEO after AI for the broader priorities; here I focus only on the hands-on work inside Screaming Frog.
Which Screaming Frog version introduced the AI features?
AI support did not arrive in a single release; it came step by step. At first you could call ChatGPT through custom JavaScript snippets. Then direct API integration followed. That is why older tutorials and the current menus can look quite different.
- Version 21.0 (November 2024): According to the official release notes, it brought direct connections to OpenAI, Gemini and Ollama, the "Config > API Access > AI" menu, up to 100 custom prompts and a new AI tab.
- June 2025, version 22.0: Anthropic (Claude) integration, semantic similarity analysis based on LLM embeddings and the content cluster diagram arrived.
- Version 24.0 (May 2026): An MCP server arrived that lets AI assistants such as Claude drive the crawler.
In short, if you run anything older than version 21, most of the menus in this guide will not appear. Update first.
After updating, also review your saved configuration profiles. Newer releases sometimes move a setting or change a default value. For instance, you can keep an old ChatGPT setup based on custom JavaScript, but the direct integration usually does the same job faster. So validate the switch with a small test crawl and compare both methods side by side.
Do you need a paid licence for Screaming Frog AI?
Yes, in practice you do. API connections and saved configurations only work in the licensed version. The free version also caps each crawl at 500 URLs. Therefore, serious AI work on a real website requires a licence.
However, the licence is not the only cost. When you connect OpenAI, Gemini or Anthropic, you pay that provider separately for tokens on every request. A ChatGPT Plus subscription does not cover API usage; you open a separate API account and add credit. Ollama, by contrast, runs the model on your own computer, so it produces no token bill, but it needs capable hardware.
With a local model the cost shifts from money to time. A small model on a laptop can spend several seconds per page, and across tens of thousands of URLs that adds up to hours. So I mostly choose local models for sensitive projects and early experiments. On large public sites, a cloud provider usually delivers better speed and quality.
I always explain this split to clients up front. Because sending a long prompt to every page of a 50,000 page site can easily produce an API bill larger than the licence itself.
How do you set up the Screaming Frog AI connection?
Setup takes a few minutes. I combined the flow from the official configuration guide with my own routine.
- Create an API account with your provider (OpenAI, Gemini or Anthropic), add credit and generate an API key.
- In the SEO Spider, open Config > API Access > AI and choose the provider.
- Paste the key into the account information field and confirm the connection with "Connect".
- In the prompt configuration area, use "Add from Library" to add a preset or write your own prompt.
- Choose which content the prompt should use (page text, HTML or a custom extraction) and which model to call.
- Test on a small crawl of 20 to 50 URLs first, then open it up to the whole site.
I also recommend saving the setup as a configuration profile. That way you never rebuild the same settings from scratch for a new project.
The most common setup error I see is an API account without credit. Your key looks fine, yet requests fail and the AI tab stays empty. The second common problem is rate limits: new API accounts may allow only a low number of requests per minute. In that case, lowering crawl speed or asking the provider for a higher limit solves it.
Which AI provider should you choose?
The right provider depends on your budget, your data privacy needs and the task. The table below summarises my own evaluation criteria. I left prices out on purpose, because model pricing changes often; check the provider's own pricing page instead.
| Provider | Strength | Watch out for | Good fit |
|---|---|---|---|
| OpenAI | Wide model range, embeddings and image generation support | Token cost, data goes to the API | Alt text, classification, embeddings |
| Gemini | Handles long content well, embeddings support | Quotas and regional settings | Text heavy pages, summaries |
| Anthropic (Claude) | Follows instructions closely, consistent text output | You may need another provider for embeddings | Meta description drafts, content review |
| Ollama (local) | Data stays on your machine, no token bill | Needs hardware, slow on big sites | Confidential projects, test crawls |
In practice I start most projects with a cloud provider. That said, for clients whose contracts forbid sending data outside, a local model through Ollama is the better choice.
Model choice matters as much as provider choice. A small, fast model from the same provider handles simple tasks such as classification well, and it costs far less. For meta description drafts, where text quality matters, a stronger model makes sense. In short, assigning a model per prompt keeps the budget balanced.
What is the difference between custom JavaScript and the direct AI integration?
Both methods do the same job, but setup and cost differ. The custom JavaScript route works through Config > Custom > Custom JavaScript and uses Screaming Frog's ready made "(ChatGPT)" snippets. For this route you must enable JavaScript rendering under Config > Spider > Rendering.
The direct integration arrived with version 21 and removed the need to edit code. The release notes state that JavaScript rendering mode is not required and that data can come back in any crawl mode. As a result, crawls run faster and use fewer resources.
So is the JavaScript route obsolete? Not entirely. When you need to work on the page as the browser builds it, for example product descriptions loaded on the client side, running prompts on the rendered DOM still helps. Still, for beginners the direct integration is the safer starting point.
Whichever method you pick, do not run both for the same task in the same crawl. Otherwise you send two API requests per page and pay twice.
How do you use the prompt library (Add from Library)?
The "Add from Library" button in the prompt configuration screen lists example prompts prepared by Screaming Frog. The version 21 notes describe them as half a dozen prompts for inspiration; alt text generation, language detection and embeddings extraction are among them.
I never use presets as they are; I treat them as templates. For instance, I add a language instruction, a character limit and a rule not to repeat the brand name to the alt text prompt. That way the output comes close to publishable.
You can also save your own prompts to your library. If you work as an agency or a team, this helps a lot: everyone runs the same approved prompt, and results stay comparable across projects.
When I write a prompt, I make sure it has four parts: role, task, constraints and output format. For example: "You are an e-commerce editor; describe the main product on the page in one sentence; stay under 120 characters; return plain text only." A clear output format matters most, because you will filter the results in a table. Answers with greetings, explanations or bullet points break that filtering.
How do you generate alt text in bulk with Screaming Frog AI?
In my experience alt text generation gives the fastest return. On large e-commerce sites, thousands of images have either empty alt attributes or file names. AI closes that gap quickly at draft level.
The flow works like this: first you filter pages with missing alt text, then you apply an image description prompt only to that segment. Version 22 added the option to run prompts only against URLs that match a specific segment. Because of that, you do not waste tokens on pages that already have alt text.
- Add the language, a length limit (for example under 125 characters) and the rule "describe what is in the image, no advertising language" to the prompt.
- Ask it to take details such as product name and colour from the page title.
- Export the output, have an editor review it, then upload it to your CMS.
Still, let me stress this: AI sometimes misidentifies the product in an image. Therefore, I do not recommend skipping human review.
Also remember what alt text is for. It exists first for visually impaired users and screen readers; the SEO benefit is a side effect. So alt text stuffed with keywords hurts both accessibility and quality. That is why the line "describe the image, do not write marketing copy" is critical in your prompt.
How do you draft meta titles and descriptions with AI?
Missing, duplicate or overly long meta descriptions are a classic Screaming Frog finding. With AI you do not only report the finding; you get a draft fix in the same crawl.
When writing the prompt, choose page text as input and lock the output into a strict pattern. For example: "Write a plain description under 155 characters that states the main topic of the page in the first sentence." Then check the results for length and keyword fit. I never leave keyword choice to the model; I match it to the target from a keyword mapping exercise.
On the other hand, publishing these drafts in bulk is risky. Google also rewrites meta descriptions from page content quite often. So prioritise pages with high impressions and low click through rate. For single page checks, you can also use the SEO checker tool.
How do you classify search intent and content type?
On large sites the hardest job is seeing which page serves which intent. With an AI prompt you can assign every page one fixed label such as "informational", "comparison", "transactional" or "brand".
The critical point is not to give the model room for free text. List only the allowed labels in the prompt and add "return nothing else". That way the output becomes a filterable column. Then you join that column with URL structure, traffic and conversion data.
For example, if a service category has no transactional page at all, that is a direct content gap. If five blog posts serve the same intent, that points to possible cannibalisation. I recommend handling such findings together with your category plan; I cover that in detail in category structure for large websites.
You can also use classification for content type. With labels like "guide", "list", "product", "category" and "news", you get the content mix of the whole site. That table is a strong starting point for showing which content types bring traffic and conversions. However, do not put the labels in a report before checking their accuracy on a small sample.
How do embeddings and semantic similarity analysis work?
An embedding turns the meaning of a text into a numeric vector. If the vectors of two pages sit close together, they cover the same topic even when their wording differs. Screaming Frog built this analysis into the tool with version 22.
According to the official version 22 release notes, the flow is: connect an AI provider, add the embeddings prompt from the library and enable the feature under Config > Content > Embeddings. After the crawl finishes you must run crawl analysis; otherwise the filters stay empty.
You see the results in the Content tab. The "Semantically Similar" filter shows pages that are very close in meaning, and the "Low Relevance Content" filter shows pages that drift away from the site's main topic. Together these two filters move duplicate content audits from word matching to meaning matching.
In short, embeddings answer the question "Why are these two articles competing for the same query?" with data.
One warning: embedding quality depends on how clean the input text is. If menus, footers and cookie notices repeat on every page, pages can look artificially similar. Therefore, defining the content area correctly, so that only the main text goes into the analysis, greatly improves reliability.
What does the Content Cluster Diagram show?
The content cluster diagram plots embeddings data on a two dimensional map. You open it via Visualisations > Content Cluster Diagram. Pages close in meaning sit close on the map, and different topics form separate clusters.
This visual helps me a lot when I explain site architecture to clients. For example, seeing blog posts float on a separate island away from service pages makes the need for internal links obvious at a glance. Also, isolated dots outside the clusters are often old campaign pages or off topic content.
Use the diagram as a decision tool. List pages that sit very close together as merge candidates, and distant ones as candidates for removal or repositioning. Then strengthen the connections between clusters with an internal linking strategy.
How do you find duplicate and cannibalising content with Screaming Frog AI?
The classic method relies on text similarity filters such as "Near Duplicates". That method catches nearly identical sentences but misses pages that cover the same topic in different words. The embeddings analysis in Screaming Frog AI closes exactly that gap.
In my own audits I follow this order:
- I export the pairs with the highest similarity scores from the Semantically Similar filter.
- For each pair, I check in Search Console whether both URLs get impressions for the same queries.
- I merge pairs that compete for the same query and 301 redirect to the winning URL.
- I separate pairs that serve different intents and sharpen their titles and introductions.
Note that high similarity alone is not a problem. Two colour variants of the same product will naturally look similar. So base the decision on search data, not on the score.
After a merge, update internal links too. Pages that link to the old URL should not send users through a redirect chain; point those links straight to the winning page. Then check in the next crawl that the pair no longer appears. That way you confirm with data that the fix worked.
How do you combine custom extraction with AI?
According to the version 21 notes, a prompt can take page text, HTML or a custom extraction as input. Custom extraction is, in my view, the least known and most efficient option. Because you send the model only the piece it needs instead of the whole page.
For example, you can pull only the product description field with XPath and ask the model "Does this description include material, size and usage information?". As a result, token cost drops and the answer gets sharper. Noise such as menus, footers and sidebars never reaches the model.
You can check structured data the same way. Extract the JSON-LD block from each page, ask the model which properties are missing, and then prepare the fix with the schema generator. Still, the final word on structured data validation belongs to Google's own testing tool.
Another benefit of custom extraction is consistency. When you send the whole page, the model sometimes mistakes a campaign banner in the menu for the main content. Sending only the target area removes that confusion. For instance, you can extract the SEO text on category pages and ask "Is this text about the category, or generic filler?".
How do you control costs when using Screaming Frog AI?
The key to cost control is how many URLs you send the prompt to and how much text you include. Token charges apply to both the input and the response. Therefore, long pages and long answers inflate the bill quickly.
- Apply prompts only to HTML pages with a 200 status that are indexable.
- Target only the relevant section, for example a /product/ folder, with a segment filter.
- Send only the necessary field through custom extraction instead of the whole page.
- Ask for short, fixed format answers such as "one word" or "at most 150 characters".
- Set a monthly spending limit in the provider dashboard.
- Lower crawl speed under Config > Speed so you do not hit API rate limits.
With these steps you run the same analysis at a much lower cost. Moreover, a small test crawl lets you measure the average token use per page and estimate a budget for the whole site.
What should you watch out for regarding data privacy?
When you connect a cloud provider, page content travels to that provider's servers. If you crawl a public website, that is usually fine. However, the situation changes when you crawl password protected areas, staging environments or customer portals.
That is why I follow three rules. First, I never run AI prompts on areas behind a login. Second, I exclude pages that may hold personal data, such as orders, accounts and form results, from the crawl. Third, for clients whose contracts restrict data sharing, I switch to a local model such as Ollama.
Also read the provider's API data usage policy. Terms such as whether your data may train models can differ for business accounts. In short, do not start a large crawl before you know where your API data goes and how long the provider keeps it.
API key security is a separate topic. Do not share keys in team chats; give each user their own key or use the provider's project based keys. Revoke a key immediately when someone leaves or when it leaks. Also, when you export and share a configuration profile, check that no key remains inside it.
How does the MCP server change the way you use Screaming Frog?
The MCP server in version 24 reverses the direction. The crawler no longer only calls AI; an AI assistant can now drive the crawler. According to the official version 24 release notes, you can start crawls, analyse data and export it from within assistants such as Claude and LM Studio.
For example, you can tell the assistant "crawl the site, list pages that return 404 and still receive internal links, then summarise". That way repetitive reporting runs in natural language. Still, the Screaming Frog team clearly states that this is not a replacement for an experienced SEO professional and that not every feature is supported yet.
My view points the same way. MCP speeds up routine work, but prioritising findings, tying them to business goals and writing the right tickets for developers remain human work.
If you want to try MCP, start with a small site. Watch step by step which command the assistant runs and which data it exports. Then compare the results with a report you built manually. That way you learn where the assistant is strong and which points you still need to check yourself. This pilot also gives you a solid base for a safe usage guide inside your team.
How far should you trust Screaming Frog AI output?
AI output is a draft, not a decision. Models sometimes invent information that is not on the page, and sometimes they answer outside the allowed label list. For that reason I sample part of the output by hand on every project.
A practical rule: check 30 to 50 random rows after the first crawl. If the error rate is high, narrow the prompt, add examples or switch to a stronger model. Then test again. Also keep in mind that the same prompt may return slightly different results on different days.
Adding one or two example answers to the prompt improves consistency. When you show the model what a correct output looks like, format errors drop noticeably. In addition, if the provider allows it, a lower temperature makes the output more predictable.
Use AI as an assistant for content quality reviews as well. Asking a model to score expertise, experience and trust signals is an interesting start, but the final judgement should stay with an editor. You can find the framework behind those criteria in my E-E-A-T guide.
How do you turn a Screaming Frog AI audit into a workflow?
A one off crawl produces interesting results, but the real value comes from a regular workflow. The flow my team and I use looks roughly like this: a planned monthly crawl, the same prompt set, a comparison report and a priority list.
First, we save the setup as a profile. Then a planned crawl runs the same settings every month. Next, we compare the AI tab with the previous month and list newly missing alt text, new cannibalising pages and off topic content.
Finally, we hand findings to developers and the content team as separate tasks. This comparison becomes even more valuable during migrations and redesigns; I collected the steps for that period in my website migration SEO checklist. If you want to set up this kind of audit for your own site, my team and I can plan it with you as part of our SEO consulting service.




