What are LLM bots?
LLM bots are the crawlers of the companies that build large language models and AI assistants. They visit public websites the way search engine crawlers do, and they introduce themselves under their own name.
CloudFilt recognizes the main ones:
- GPTBot, the crawler of OpenAI, behind ChatGPT
- ClaudeBot, the crawler of Anthropic, behind Claude
- PerplexityBot, the crawler of the Perplexity answer engine
- Gemini, the AI crawling of Google
- DeepSeekBot, GrokBot and CohereBot
What are they used for?
An LLM bot reads your pages for two reasons.
Answering questions in real time
When someone asks an AI assistant about a product, a service or a topic you cover, the assistant can fetch your pages to build its answer, and often cites them with a link. More and more visitors discover websites this way, the same way they used to through a search engine.
Training and updating models
AI companies also collect public pages to train their models and keep their knowledge up to date. What the model learns about your brand, your products and your expertise depends on what it could read.
Why CloudFilt lets them through by default
On paper, an LLM bot looks like everything CloudFilt is built to stop: it reads many pages in a row, runs in a datacenter, and never executes JavaScript. Without special care, the scraping, headless browser and hosting network checks would block it on its first visit.
Blocking it by default would make your website disappear from the answers of ChatGPT, Claude, Perplexity or Gemini, without you ever deciding it. That is why known LLM bots are allowed by default:
- Their identity is verified. CloudFilt checks where the request really comes from. A bot that only copies the name of GPTBot in its User-Agent gets no free pass: your bans and your flood limits still apply to it.
- A verified crawler is never banned for good. If a traffic spike got it banned, the ban is lifted as soon as its identity is confirmed, so you do not vanish from AI answers because of one busy afternoon.
- Your own rules still come first. Your access lists and your country rules apply to LLM bots like to any other visitor.
Why you may want to block them
Being read by AI assistants is not a good deal for every website. You may prefer to keep LLM bots out when:
- your content is your product: paid articles, research, databases, price catalogs;
- you do not want your texts or your images used to train models without an agreement;
- their crawling costs you more server resources than it brings you visitors;
- you negotiate the use of your content with AI companies and want to control access in the meantime.
A robots.txt file only asks crawlers to stay away. CloudFilt enforces your choice on every request.
Block or limit LLM bots in one click
In your dashboard, open the settings of your website, then the Protection tab. The AI crawlers line gives you two choices:
- Block: every known LLM bot is refused, on every request.
- Rate limit: LLM bots stay allowed, but each one is capped to the number of requests per minute you set. Above it, its requests are refused until the next minute.
The change applies right away, and you can save it for all your websites at once. Nothing is permanent: LLM bots are never banned, so switching the option back lets them in again on their next visit.
Refused requests show up on the Bots page of your dashboard, next to the rest of your traffic.