May AI SOLUTIONS

Common questions · Websites

Should I block AI bots from crawling my website?

For nearly every small business, no. Which crawlers are which, what blocking really costs you, and how to check what your own site allows today.

A heavy wooden bar latch across the planks of an old barn door, held shut by a single round peg, the timber weathered grey and brown.

Somebody has told you that AI companies are helping themselves to your website, and that you can stop them. The first part is true and the second part is mostly true. Whether you should is a different question, and for nearly every small business the answer is no.

Blocking those crawlers is the surest way to guarantee you are never the business ChatGPT names.

Two different visits, and they are not the same bot

This is where most of the advice goes wrong, so it is worth separating properly.

The training visit. A crawler reads your pages and the text goes into the pile a model is built from, months before anybody asks a question. Slow, indirect, and honestly not worth much to you either way.

The answering visit. Somebody asks an assistant for a plumber in your town. It runs a few searches and fetches pages right then, in the seconds before it writes its answer. That fetch is also a bot — and it is the one deciding whether your name appears. I have written about what happens in those seconds before.

Almost every “block the AI scrapers” post you will find was written about the first visit. The instructions get pasted in and shut the door on both.

The names on the door

The file doing the blocking is called robots.txt, and yours already exists. Type your own address followed by /robots.txt and you can read it. It is plain text, it names crawlers, and it tells each one what it may look at. Every crawler has its own name, and a rule aimed at one name does nothing to the rest:

  • GPTBot — OpenAI’s training crawler.
  • OAI-SearchBot — the one that indexes pages so ChatGPT can find them at all.
  • ChatGPT-User — the fetch that happens because a person’s question sent it to your page.
  • Google-Extended — not really a crawler, more a switch about whether Google may use your content in Gemini and its AI products. Google’s own documentation says it does not change whether you appear in Google Search.
  • ClaudeBot, PerplexityBot, and others. The list gets longer every few months.

Which is why the copied-and-pasted block lists doing the rounds tend to be half stale and half wrong: they name last year’s bots, they miss the ones that matter, and they rarely distinguish the crawler that trains from the crawler that answers.

Two more things about that file. It is a request, not a lock — well-behaved companies honour it and badly behaved ones never asked. And it is public, so it is no way to hide anything. Anything that genuinely must stay private belongs behind a login, not behind a polite note.

What blocking actually costs you

Plainly: if the answering crawler cannot read your page, you are not in the answer. The assistant does not leave a gap where you would have been. It names the competitor two towns over, in a complete and confident sentence, and your customer never learns you existed.

There is no notification when this happens. It is the same silence as the visitor who left because your site was slow — no missed call, no half-filled form, nothing in your figures to point at.

The traffic argument, and why it does not transfer to you

The businesses blocking hardest are publishers — newspapers, magazines, recipe sites. Their income is the page view. An assistant that answers the question without sending anybody along takes money straight out of the till, and they have a real grievance.

Your maths runs the other way. You do not sell page views. You sell roofs, or appointments, or brake jobs. Being named in an answer ends in a phone call, a map pin, or somebody typing your name into Google an hour later. Somebody who never hears your name does none of those things. You are not defending a toll gate; you are trying to get mentioned.

Go and look at what your site says right now

The accidental block is far more common than the deliberate one, and that is the real reason to check.

Some hosting companies and security services now block AI crawlers by default, or put it behind a single switch that somebody may have flipped without thinking. SEO plugins add rules. “Block bad bots” plugins add more. Plenty of small businesses are turning these crawlers away today having never made a decision about it.

Open yourbusiness.com/robots.txt and look for Disallow: / sitting under any of the names above, or under User-agent: *. If you find one you did not choose, whoever looks after your site can take those lines out in about five minutes.

Worth being straight about, because the opposite gets sold hard. An open door is not a customer walking through it. You still have to be findable in ordinary search, say the same name, hours and phone number everywhere you appear, keep reviews coming, and state what you do in plain sentences a machine can lift — the working list is already written out.

And nobody can promise you a place in any assistant’s answer. Ask the same question twice and it changes. Anyone guaranteeing you a spot is selling something they do not control.

The short version

Blocking AI crawlers protects a business whose product is the page itself. Yours is not that business. The bot that trains a model and the bot that fetches your page mid-answer are different bots with different names, and the blanket instructions circulating online stop both. Read your own robots.txt this week — the block you need to worry about is the one you never chose.

The free site check reads your pages the way one of these machines does and tells you what it could and could not find in them. It does not read your robots.txt, so that one is on you, and it takes a minute. Or tell me your trade and your town and I will look at both.

If this is the question you came with and you would rather just ask a person, that is what the form is for. I read them myself.

Tell me what you need