Illustration of a website controlling which AI crawlers are allowed or blocked

Should You Block AI Crawlers From Your Website?

I do not think there is one correct answer to whether you should block AI crawlers.

That sounds less exciting than telling everybody to block everything or allow everything, but it is the useful answer.

Your decision depends on what you are trying to accomplish.

If you publish original work and do not want certain crawlers collecting it for model training, that is one concern. If you run a business and want people using AI search to discover your public pages, that is another. They can involve different crawlers and different rules.

Decide what you actually want first

Before touching robots.txt, I would answer a few questions.

Do I want this public content discoverable through AI search? Do I want to restrict crawlers associated with model training? Are there private areas that should not be public at all? Am I trying to solve a copyright concern, a bandwidth problem or a search visibility problem?

Those are not the same problem, so they should not automatically get the same configuration.

OpenAI separates search access from training access

OpenAI’s current publisher documentation tells site owners who want their content discoverable in ChatGPT search not to block OAI-SearchBot. GPTBot is a separate crawler associated with content that may be used to improve models.

That separation means a publisher can make a more specific decision than “OpenAI yes” or “OpenAI no.”

Read OpenAI’s current publisher guidance before changing the rules.

Google’s controls are another example

Google says its Google-Extended token controls certain uses of crawled content for Gemini models and grounding. Google specifically says Google-Extended does not affect inclusion or ranking in Google Search.

Google explains Google-Extended in its crawler documentation.

The details matter. A crawler name in robots.txt does not tell you its purpose by itself.

Robots.txt does not protect private information

This is important enough to say plainly.

If something genuinely needs to be private, do not rely on robots.txt to protect it. Use authentication and proper access controls.

Robots.txt is a crawling instruction for compliant bots. The file itself is public.

What I am doing on my own sites

My preference is to make these decisions crawler by crawler based on what each one does rather than copy a blanket rule from somebody else’s site.

I want public Marketur content discoverable. I also want to know which systems are accessing it and why. As the ecosystem changes, I can revisit those choices.

That is the reason I built a free AI Readability Check. It reads the current configuration and separates the different kinds of access instead of pretending “AI crawler” is one category.

Check before you change anything

If you are not sure what your website currently allows, start there.

Check your website’s AI crawler access on Marketur. Once you know what the existing rules actually do, deciding whether to change them gets a lot easier.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top

More

☀️ Light Mode
📰 Latest Posts
Loading...
📊 Community Stats
Loading...
🟢 Online Now
Loading...
👋 New Members
Loading...
👥 Popular Groups
Loading...