For a while the advice around AI crawlers was basically one sentence: block the AI bots.
I understand why. Site owners saw AI companies crawling the web and wanted control over how their content was used. The problem is that the crawler landscape is more specific now, and treating every bot as if it does the same job can create a completely different result than you intended.
The clearest example is OpenAI’s GPTBot and OAI-SearchBot.
GPTBot and OAI-SearchBot are not interchangeable
OpenAI separates its search crawler from the crawler associated with potential model training.
OAI-SearchBot is used for search discovery. OpenAI tells publishers that if they want their content discoverable and eligible to appear in ChatGPT summaries and snippets, they should make sure OAI-SearchBot is not blocked.
GPTBot is a separate control related to content that may be used to improve generative AI models.
You can read OpenAI’s current guidance in its publisher and developer documentation.
Why I care about the distinction
If my goal is to prevent a particular kind of model-training crawl, I want to make that decision deliberately.
But I do not automatically want that decision to prevent somebody using ChatGPT from discovering one of my public articles, tools or products.
Those are two different business decisions.
This is similar to how I think about WordPress in general. I would rather understand the switch I am flipping than install somebody else’s configuration and hope their goals match mine.
Robots.txt is not a universal privacy switch
A robots.txt rule tells compliant crawlers which paths they are allowed to crawl. It does not make public information private, and it is not a replacement for authentication or access control.
OpenAI also notes that a blocked page can sometimes still have its URL and title surfaced when the URL is discovered through other sources. If you actually need a page excluded from search results, that becomes a different technical question.
Other companies have their own controls
OpenAI is not the only company with crawler-specific controls. Google documents Google-Extended as a separate robots.txt token that controls certain uses of crawled content for Gemini models and grounding while not controlling normal Google Search inclusion.
That is another example of why a giant “block AI” rule can be too blunt for what a site owner is actually trying to accomplish.
Google documents Google-Extended with its other crawler controls.
See what your own site is doing
I built the Marketur AI Readability Check because I wanted this information in one place.
Give it a URL and it checks the site’s crawler rules along with the page itself. The goal is not to tell everybody they must allow or block a particular company. The goal is to show you what your current configuration does so you can decide what you actually want.
If you changed your robots.txt during the first wave of AI crawler blocking, this is worth checking again now.
