Try for free Log in

Known LLM bots

GPTBot, ClaudeBot, PerplexityBot, Gemini: the crawlers of AI assistants read your pages to answer their users. CloudFilt lets them through by default, and you can cap or block them in one click.

Free up to 3,000 requests per month and per site. No credit card required.

Verified crawlers Your call

Known LLM bots

Reads
Public pages and content
Your controls
Verified identity, Block in one click, Rate limit

How CloudFilt handles known LLM bots

What are LLM bots?

LLM bots are the crawlers of the companies that build large language models and AI assistants. They visit public websites the way search engine crawlers do, and they introduce themselves under their own name.

CloudFilt recognizes the main ones:

  • GPTBot, the crawler of OpenAI, behind ChatGPT
  • ClaudeBot, the crawler of Anthropic, behind Claude
  • PerplexityBot, the crawler of the Perplexity answer engine
  • Gemini, the AI crawling of Google
  • DeepSeekBot, GrokBot and CohereBot

What are they used for?

An LLM bot reads your pages for two reasons.

Answering questions in real time

When someone asks an AI assistant about a product, a service or a topic you cover, the assistant can fetch your pages to build its answer, and often cites them with a link. More and more visitors discover websites this way, the same way they used to through a search engine.

Training and updating models

AI companies also collect public pages to train their models and keep their knowledge up to date. What the model learns about your brand, your products and your expertise depends on what it could read.

Why CloudFilt lets them through by default

On paper, an LLM bot looks like everything CloudFilt is built to stop: it reads many pages in a row, runs in a datacenter, and never executes JavaScript. Without special care, the scraping, headless browser and hosting network checks would block it on its first visit.

Blocking it by default would make your website disappear from the answers of ChatGPT, Claude, Perplexity or Gemini, without you ever deciding it. That is why known LLM bots are allowed by default:

  • Their identity is verified. CloudFilt checks where the request really comes from. A bot that only copies the name of GPTBot in its User-Agent gets no free pass: your bans and your flood limits still apply to it.
  • A verified crawler is never banned for good. If a traffic spike got it banned, the ban is lifted as soon as its identity is confirmed, so you do not vanish from AI answers because of one busy afternoon.
  • Your own rules still come first. Your access lists and your country rules apply to LLM bots like to any other visitor.

Control LLM bots for free

Why you may want to block them

Being read by AI assistants is not a good deal for every website. You may prefer to keep LLM bots out when:

  • your content is your product: paid articles, research, databases, price catalogs;
  • you do not want your texts or your images used to train models without an agreement;
  • their crawling costs you more server resources than it brings you visitors;
  • you negotiate the use of your content with AI companies and want to control access in the meantime.

A robots.txt file only asks crawlers to stay away. CloudFilt enforces your choice on every request.

Block or limit LLM bots in one click

In your dashboard, open the settings of your website, then the Protection tab. The AI crawlers line gives you two choices:

  1. Block: every known LLM bot is refused, on every request.
  2. Rate limit: LLM bots stay allowed, but each one is capped to the number of requests per minute you set. Above it, its requests are refused until the next minute.

The change applies right away, and you can save it for all your websites at once. Nothing is permanent: LLM bots are never banned, so switching the option back lets them in again on their next visit.

Refused requests show up on the Bots page of your dashboard, next to the rest of your traffic.

Control LLM bots for free

From request to verdict

Each request is scored from the front and the back end, then allowed, challenged or blocked before it reaches your server.

  1. Visitor

    A browser, a script or a bot sends a request to your website or API.

  2. Nearest point of presence

    The CDN WAF, or your plugin, hands the request to the closest point of presence.

  3. Signals

    IP reputation, behaviour, rate, country and your own rules, weighed in real time.

  4. Verdict

    • Allowed to your site
    • Challenge captcha first
    • Blocked never reaches you

Other threats CloudFilt stops

All solutions