TechnicalStrategyGEO

Should You Block AI Crawlers? The Case Both Ways

The instinct to block AI crawlers is understandable. Models trained on your work compete with it, and answers built from your content can satisfy the reader without ever sending them to you. But blocking is not a neutral defensive move. It removes you from the surfaces where discovery is increasingly happening. The right answer depends on how you make money, and it is genuinely different for different businesses.

First, know which control does what

These get conflated, and the consequences differ sharply.

ControlWhat it stopsWhat it does not stop
GPTBot disallowOpenAI crawling for trainingChatGPT browsing live for an answer
OAI-SearchBot disallowChatGPT search retrievalTraining crawls
Google-Extended disallowUse for Gemini training and groundingGooglebot, and your classic rankings
Search Console AI toggleAppearing in AI Overviews and AI ModeAppearing in normal search results
Googlebot disallowEverythingNothing. Do not do this

The mechanics of each user-agent, and the exact syntax, are covered in robots.txt for AI crawlers. The point here is strategic: “block AI” is not one decision. It is at least four, and you can take different positions on each.

Google’s Search Console opt-out, which began testing in June 2026 with a subset of UK site owners after intervention from the UK Competition and Markets Authority, is the bluntest of them. Google’s own framing is unambiguous: sites that opt out “will not receive traffic or impressions from our generative AI features.” It is not applied as a ranking signal for core web search.

The case for blocking

Your content is the product. If people pay for access to what you publish, letting a model absorb and resummarize it is straightforward substitution. Publishers, research firms, and subscription media have a real argument here, and roughly a third of publishers surveyed said they intended to block Google’s generative features.

You have licensing leverage. If your archive is valuable enough that AI companies would pay for it, giving it away free removes any basis for negotiation. Blocking first and licensing second is a coherent commercial strategy for a large enough corpus.

The material is genuinely proprietary. Original datasets, methodologies, and client-confidential material should not be sitting in a training corpus regardless of the traffic tradeoff.

Crawl cost is real. For very large sites, aggressive AI crawling consumes meaningful bandwidth and origin capacity, which is what drove the emergence of per-crawl pricing and crawler-management products at the CDN layer.

The case against blocking

Discovery is moving into the answers. If a buyer asks an assistant to shortlist vendors and you are not retrievable, you are not on the list. There is no second page to be on. For a business that needs to be found, this is the whole argument.

You lose the brand impression, not just the click. Most AI answers are zero-click. The value is frequently the mention itself, being named as the credible option. Opting out forfeits that exposure while your competitors keep it.

Blocking training does not reclaim what is already learned. Models trained before your block still hold what they absorbed. You pay the future cost without recovering the past one.

Partial blocking often backfires quietly. Disallowing a search-retrieval bot while leaving the training bot open is a common misconfiguration that gets the worst of both: your content still feeds the model, and you still cannot be cited.

How to decide

Decision tree for whether to block AI crawlers A decision tree starting from how the business makes money. If revenue comes from content access, the next question is whether you have licensing leverage: with leverage, block training bots and negotiate; without it, block only proprietary paths. If revenue comes from being hired, the next question is whether absence from AI answers would cost pipeline: if yes, allow retrieval bots and decide on training bots separately, which is the recommended default for most service businesses. A footnote notes that crawl cost should be handled by rate limiting at the CDN, never a blanket block. Blocking is at least four decisions, not one Work it through per bot and per path, starting from how you make money How do you make money? start here Selling access to content publisher, research, subscription Selling services or products you need to be discovered Licensing leverage? is the archive big enough to sell? Would absence cost pipeline? for most, yes Block training bots, negotiate leverage: yes Block proprietary paths only leverage: no Allow retrieval bots Decide on training bots separately the default for most businesses Whichever branch you land on: if the real concern is crawl load, rate-limit at the CDN. Never blanket-block for cost.

Scroll the diagram sideways to see all of it.

Note: Google's Search Console opt-out for AI Overviews and AI Mode began testing June 2026 with a subset of UK site owners, following intervention by the UK Competition and Markets Authority. Google states opted-out sites receive no traffic and no impressions from generative AI features, and that it is not a ranking signal for core web search.

Work through it in this order.

  1. Is your revenue from content access or from being hired? Content access leans toward blocking. Being hired leans strongly against it.
  2. Would being absent from AI answers cost you pipeline? For most service businesses, professional firms, local businesses and SaaS, yes, materially.
  3. Do you have licensing leverage? Only a large, distinctive corpus does. If not, blocking buys protection you cannot monetize.
  4. Is the concern actually crawl load? Then rate-limit at the CDN rather than blocking outright. That solves the cost problem without the visibility cost.
  5. Are specific sections the real issue? Block those paths and leave the rest open.

The middle path most businesses should take

Blanket positions are rarely right. A defensible default for a business that sells services rather than content:

  • Allow retrieval bots. OAI-SearchBot, PerplexityBot and equivalents are how you get cited. Let them in.
  • Decide on training bots separately. Blocking GPTBot and Google-Extended while allowing retrieval is a legitimate stance: be quotable now, without contributing to the next model.
  • Protect specific paths, not the whole site. Disallow gated resources, client work, and proprietary datasets by path.
  • Rate-limit rather than block for load. Handle crawl cost as an infrastructure problem.
  • Stay out of the Search Console AI opt-out unless you have concluded you want no presence in Google’s AI surfaces at all.

Note the tension in the second point: allowing retrieval while blocking training is coherent, but some engines ground answers using the same infrastructure they train on, so the separation is cleaner in policy than in practice. Verify behaviour rather than assuming it.

Where this leaves you

For publishers, blocking can be a rational defence of the product. For everyone whose website exists to win work, blocking AI crawlers mostly means volunteering to be invisible in the channel that is growing fastest. Decide per bot, per path, and on the basis of how you actually make money, not on how the training debate makes you feel.

If you want your crawler policy reviewed against what your competitors are allowing, and what it is currently costing you in citations, that is part of our GEO service.

Frequently asked questions

Does blocking AI crawlers hurt my Google rankings?

Blocking Google-Extended does not affect core web search rankings, and Google says its Search Console opt-out for AI features is not a ranking signal outside those features. Blocking Googlebot itself would remove you from search entirely, which is why the distinction matters.

What do I lose by opting out of AI Overviews and AI Mode?

Google states that sites which opt out receive no traffic and no impressions from its generative AI features. You lose the citation, the referral, and the brand exposure inside answers, while remaining in classic search results.

Is blocking AI crawlers reversible?

The access control is, but the effect is not immediate or symmetrical. Removing a block restores crawling, yet content already absorbed into training corpora does not come back out, and rebuilding retrieval presence takes time.

Who should actually block AI crawlers?

Publishers whose product is the content itself, sites under licensing agreements, and businesses with paywalled or proprietary material. Service businesses that need to be discovered almost always lose more than they protect.

OK

Olga Kunger

Founder & Lead Strategist, Ambeltek

Olga leads Ambeltek's web development, AI SEO, and GEO work — helping brands rank on Google and get cited by AI engines. More about Olga →

Ready to get found on Google and in AI?

Tell us about your project and we’ll send a free visibility audit plus a tailored proposal, usually within one business day.

Start a project