Faro AI Signals·August 15, 2026

Citation Gap Hits 46× as Crawler Blocks Reach Half the Web

New data shows brand citation rates vary by 46× across AI platforms, while more than half of Google AI Mode-cited domains are actively blocking at least one AI crawler — a combination that makes platform-diversified visibility the defining challenge for businesses in AI search.

ShareShare

This Week's Signals

1

Brand Citation Rates Span 46× Across Platforms as Gemini Hits 1.5 Billion Users

Source: Everything PR / Leapd AI

Why it matters for your score

A study of 34,234 AI responses found ChatGPT cited brands 0.59% of the time while Perplexity sat at 13.05% and Grok at 27% — meaning the same content gets cited at wildly different rates depending on which platform a customer uses. An analysis of 680 million citations found only 11% of domains are cited by both ChatGPT and Perplexity, so optimizing for one platform actively leaves most AI traffic uncovered. Separately, brands present on four or more platforms are 2.8× more likely to appear in ChatGPT responses, per a 30 million citation analysis by Peec AI.

2

51.9% of Google AI Mode-Cited Domains Block at Least One AI Crawler

Source: HasData

Why it matters for your score

HasData's baseline of 10,894 domains found that in 10 Google AI Mode test queries, 27 of 52 cited domains block at least one AI crawler in robots.txt — yet they still get cited. The study also found that blocking GPTBot does not keep a site out of Google AI Overviews, which means blanket blocking decisions made to protect content from training crawlers may not have the search-exclusion effect owners assume. With Cloudflare's September 15 default-block policy covering 8.5% of the top web, now is the time to audit exactly which crawlers your robots.txt is blocking and why.

3

ChatGPT's AI Referral Share Falls to 62.6% as llms.txt Direct Hits Total Just 408 in 500 Million Bot Visits

Source: Limy.ai / Goodie

Why it matters for your score

Limy monitored over 500 million AI bot visits across a 90-day window and found only 408 targeted llms.txt directly — GPTBot, ClaudeBot, PerplexityBot, and Google-Extended overwhelmingly skip the file and crawl HTML instead. ChatGPT's share of B2B AI referral traffic has dropped from roughly 89% a year ago to 62.6%, with Claude at 18.5% and Gemini at 10.6%, so a strategy built around a single platform now covers far less of the AI traffic landscape than it did in 2025. Limy frames llms.txt as a Business-to-Agent routing surface rather than a citation signal, which changes how businesses should think about maintaining it.

What to do this week

  1. 1

    Run your domain through Faro's AI Readiness Scan (/tools/ai-readiness-scan) to see which AI crawlers your robots.txt currently blocks — the HasData study shows that 51.9% of Google AI Mode-cited domains block at least one crawler, so understanding your exact block profile is the first step before the Cloudflare September 15 deadline.

  2. 2

    Use Faro's llms.txt generator (/tools/llms-txt) to publish or update your llms.txt file as a machine-readable routing surface for agent traffic — Limy's data confirms crawlers skip the file when it is absent or malformed, and the framing matters: treat it as a Business-to-Agent signal, not a citation shortcut.

  3. 3

    Run a competitor intelligence check (/tools/competitor-intelligence) across at least four AI platforms — Perplexity, ChatGPT, Gemini, and Grok — since the citation rate study found brands present on four or more platforms are 2.8× more likely to appear in ChatGPT responses, and only 11% of domains appear in both ChatGPT and Perplexity citations.

Sources

← All editionsScan your site now →