Can your CDN or firewall block AI crawlers when robots.txt allows them?

Quick answer

Yes. CDN and WAF bot scores can return 403 responses before robots.txt is consulted. Perplexity documents stealth crawl disputes with Cloudflare; treat vendor guidance as a signal to audit, not a verdict. Align robots.txt allows with firewall rules, disable generic “block AI bots” toggles unless policy requires them, and re-test fetches after deploy.

What are the symptoms?

Pages rank in Google, buyers see them in browsers, yet AI tools never cite them and server logs show 403 for bot user agents. Check CDN analytics for challenged requests.

What is the audit order?

  1. Export allowed search bots from your policy doc.
  2. Read robots.txt at the live domain.
  3. Review CDN bot fight settings and custom WAF rules.
  4. Fetch top URLs with each allowed user agent.

How do values calls fit?

Some brands block training bots but allow search bots. Document the list so infra changes do not undo marketing intent. See training vs search bots.

Frequently asked questions

Should we allow PerplexityBot?

If you want Perplexity to cite your proof pages, allow its documented crawler unless legal policy says otherwise.

Will rate limiting look like a block?

Aggressive throttling can mimic a ban. Check response codes and retry-after headers.

Who owns the decision?

Marketing, legal, and infra should sign one crawler matrix yearly.

Need this applied to your brand?

Book a call with fisagency, or continue in Calculators, Prompts, and Skills.