AI Visibility

AI Crawlers Ignore llms.txt 97% of the Time So Stop Making One

Faro Editorial

August 5, 2026 · 8 min read

Browser console showing llms.txt file with zero AI crawler requests logged beside it
```html

Ahrefs analyzed 137,210 live domains in May 2026 and found that 97% of llms.txt files received zero requests, from any bot, any human, anything. Not a single fetch. If you spent time creating one last quarter, you likely spent it on a file that no AI crawler ever opened. The relationship between llms.txt and AI crawlers that the SEO community hoped for simply does not exist in the data.

That finding is not a rounding error or a sampling artifact. It holds across tens of thousands of domains, and a separate 300,000-domain study by SE Ranking reached the same conclusion from a different angle. Yet llms.txt implementation guides keep circulating, agencies keep billing for the work, and founders keep asking whether they need one. The answer right now is almost certainly no.

Here is why the file never delivered what its proponents promised, what AI bots are actually doing instead, and where to point your time and budget if AI citation visibility is the goal.

What did the Ahrefs study actually find?

The numbers are more damning than the headline statistic suggests. Ahrefs found that of roughly 38,000 domains with a valid llms.txt file, only about 1,100 received any traffic to it in May 2026. That's a 97% failure rate on the basic premise that creating the file would invite AI crawlers to read it.

Of the 3% of files that did attract requests, the composition of those requests dismantles the narrative. According to Search Engine Journal's reporting on the Ahrefs data, 96% of those requests came from bots, mostly non-AI ones; AI retrieval bots linked to ChatGPT and Perplexity made up just 1% of total requests. SEO audit tools accounted for 21% of requests. Unidentified bots accounted for 14%. Slackbot fetched llms.txt files more often than PerplexityBot did.

One more finding deserves attention: Ahrefs confirmed that AI bots never requested llms.txt on domains where the file did not already exist. Crawlers are not probing for it. They're not checking whether you have one. They don't care.

What are GPTBot and ClaudeBot actually crawling instead?

The bots crawl HTML the same way Google has for two decades. The growth numbers here are striking. From May 2024 to May 2025, GPTBot request volume grew 305% and overall AI-plus-search crawler traffic rose 18%, per Cloudflare data. All of that growth happened against standard HTML pages.

Cloudflare's traffic data shows GPTBot's share of verified bot traffic grew from 4.7% in July 2024 to 11.7% in July 2025, while ClaudeBot grew from 6% to nearly 10% over the same period. Neither operator has documented their production systems as consuming llms.txt.

Cloudflare's 2025 Year in Review found crawling for model training reached as much as 7 to 8 times search crawling volume and 32 times user-action crawling at peak. That's an enormous amount of crawl activity, and it operates entirely on raw HTML.

The official documentation confirms this. OpenAI's crawler documentation covers GPTBot, OAI-SearchBot, and ChatGPT-User and instructs site owners to manage access via robots.txt. The llms.txt file isn't mentioned as a signal or a requirement anywhere in that documentation.

Does having an llms.txt file improve AI citation frequency?

No. Not in any study that has attempted to measure it. SE Ranking analyzed 300,000 domains and found no statistically significant correlation between having an llms.txt file and AI citation frequency; removing the llms.txt variable from their XGBoost predictive model actually improved its accuracy. The variable was noise, not signal.

Even the file's adoption rate undercuts the narrative. SE Ranking's crawl of those 300,000 domains found llms.txt on only 10.13% of them. Nearly 9 out of 10 sites had not implemented it, and those sites were being cited by AI tools at rates no different from those that had.

Google's guidance is equally direct. In a section titled "mythbusting" published in late May 2026, Google told site owners that machine-readable files like llms.txt are not needed to appear in generative AI search. Google's John Mueller described llms.txt as "not done for search" and called it "a temporary crutch, perhaps to save some tokens" for AI coding tools parsing developer documentation.

If you want to understand your actual AI visibility posture rather than guessing, Faro's AI Readiness Scan runs 38 checks across six categories and returns a scored view of what AI systems can and cannot read on your site today.

Wondering whether your content is actually reaching AI systems? Run a free AI Readiness Scan and see exactly which signals are helping or hurting your AI citation chances.

How does llms.txt compare to signals that actually matter?

The table below maps the main levers marketers reach for when trying to improve AI visibility. It shows what the current evidence says about each one.

Signal What AI crawlers do with it Evidence of citation lift Who controls it
llms.txt file 97% of the time: nothing No measurable effect in two large studies Site owner
robots.txt directives GPTBot respects allow/disallow blocks Controls access; documented by OpenAI Site owner
Structured data / Schema markup Parsed during HTML crawl Strong signals for entity recognition Site owner
Page HTML clarity and heading structure Consumed directly by every major AI crawler Core input to answer generation Site owner
Third-party citations and mentions Weighted heavily in training corpora Consistent predictor of AI recommendation Earned through PR, content, authority
Brand search volume Indirect signal via indexed content about brand AI-recommended brands are 389% more likely to be Googled after Earned through awareness

Why did llms.txt spread so fast if it does nothing?

The proposal had a credible origin story. Answer.AI and fast.ai published the llms.txt specification in late 2024, positioning it as a way for sites to surface clean, context-ready content for LLM consumption. That framing resonated with SEOs trained to think in terms of structured signals: if Google needed sitemaps and schema, surely AI would need its own guidance file.

The logic was plausible. The problem is that it was never validated. No major AI lab adopted the spec in their documented crawler behavior. Traffic analysis across hundreds of millions of AI requests shows GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, and Google-Extended overwhelmingly skip the llms.txt file and go straight to crawling HTML.

As of mid-2026, none of OpenAI, Google, Anthropic, Meta, Perplexity, or Mistral has publicly stated that their production systems read or act on llms.txt. The spec spread because it was easy to implement and satisfying to tick off a checklist. Neither of those is a good reason to prioritize it.

The robots.txt comparison reveals the deeper problem

robots.txt works because every major crawler has been programmed to read it, respect it, and alter behavior based on it. OpenAI says so explicitly in their crawler docs. Google has enforced it for decades. The compliance infrastructure exists.

llms.txt has no equivalent infrastructure. No crawler has committed to reading it. No enforcement mechanism exists. It's a file that site owners write, post publicly, and then wait, while bots that were never instructed to look for it continue crawling HTML as they always have. If you want to control what AI systems can see on your domain, your robots.txt configuration is the only file with actual teeth right now.

What should you do instead of building an llms.txt file?

Focus on the same things that have always driven topical authority, with an AI-visibility lens applied to each one.

First, get your HTML structure right. AI crawlers parse the same heading hierarchy, body copy, and link signals that search crawlers do. Thin pages, buried context, and weak entity signals hurt you in both channels. Fix the underlying content quality before worrying about any auxiliary file.

Second, get your structured data right. Schema markup is consumed during the HTML crawl that AI bots are already performing. Entity clarity, FAQ schema, and organization markup all give AI systems cleaner inputs when they're deciding whether to cite or summarize your content. Faro's AI Schema Creator generates markup specifically optimized for the entity signals AI systems look for.

Third, earn third-party mentions. The SE Ranking study and broader AEO research consistently identify external citations as the strongest predictor of AI recommendation frequency. A brand mentioned in authoritative sources gets pulled into training data and retrieval contexts repeatedly. A brand with a polished llms.txt file and no external mentions gets ignored.

Fourth, measure what's actually happening. Most marketers building for AI visibility have no idea whether AI systems are citing them, how often, or in what contexts. That gap is worth closing before any implementation work begins.

Frequently Asked Questions

Will llms.txt ever start working as AI search matures?

It's possible but not guaranteed. No major AI lab has committed to supporting the spec, and the Ahrefs data from May 2026 shows adoption has not translated into crawler behavior. If OpenAI, Anthropic, or Google formally integrates llms.txt into their crawler documentation, the calculus changes. Until then, there's no evidence-based reason to prioritize it over HTML structure, schema, and earned mentions.

Does llms.txt hurt anything if I already have one?

No. A correctly formatted llms.txt file doesn't appear to penalize sites or misdirect crawlers. The cost is the opportunity cost: time spent creating and maintaining it that could go toward content quality, schema implementation, or citation-building. If you have one, keeping it live is fine. Building elaborate versions from scratch isn't worth the investment given current data.

What file actually controls whether GPTBot crawls my site?

robots.txt. OpenAI's official crawler documentation instructs site owners to use robots.txt disallow directives to block GPTBot, OAI-SearchBot, or ChatGPT-User. That's the documented, tested, and respected mechanism. llms.txt has no equivalent enforcement pathway with any major crawler operator.

If 9 out of 10 sites lack llms.txt and still get cited, what does predict AI citations?

The SE Ranking study found that topical authority signals, external mention frequency, and structured HTML clarity were stronger predictors than any auxiliary file. Brands with clear entity definitions, consistent mentions in authoritative sources, and clean content structure outperform in AI citation contexts regardless of whether they've implemented llms.txt.

Should agencies stop recommending llms.txt to clients entirely?

For most clients, yes. The implementation is low-cost but not zero-cost, and the evidence of benefit is nonexistent across two large studies. Agencies that position llms.txt as a meaningful AI visibility tactic are selling a ritual, not a result. Direct that effort toward schema audits, content clarity improvements, and citation monitoring instead.

In short

The data is unambiguous: 97% of llms.txt files received zero requests in the most comprehensive study ever run on the format. GPTBot, ClaudeBot, and every other major AI crawler crawl HTML directly and have no documented behavior tied to llms.txt. Two independent large-scale studies found no correlation between having the file and appearing in AI citations more frequently. The file spread because the logic sounded plausible and the implementation was easy, not because it worked. If AI visibility is a commercial priority for your brand, the levers are HTML structure, schema markup, earned third-party mentions, and knowing your current citation baseline. None of those require a llms.txt file.

Ready to find out what AI systems actually see when they crawl your site? Run a free AI Readiness Scan and get a scored breakdown of the signals that move AI citation frequency, not the ones that just look good on a checklist.

```

Related Reading

← Back to Blog

The Faro platform

Every tool you need to be found, understood, and chosen by AI.

Faro is building the complete infrastructure layer for AI discoverability. Scan first, then fix, monitor, and stay ahead. All from one platform.

Revenue CalculatorLive

See what poor AI readiness is costing you

Use tool →
AI Readiness ScanLive

Score your site 0–100 for AI agent visibility

Use tool →
llms.txt GeneratorLive

Tell AI exactly who you are in 60 seconds

Use tool →
AI Schema CreatorLive

Generate JSON-LD structured data for AI agents

Use tool →
Competitor IntelligenceLive

Side-by-side AI readiness scores vs competitors

Use tool →
robots.txt AnalyzerLive

See which AI crawlers your site is blocking

Use tool →
Pricing Clarity AuditorLive

Is your pricing page readable by AI agents?

Use tool →
OKF GeneratorLive

Machine-readable knowledge bundle for AI agents

Use tool →
WebMCP Readiness CheckLive

Check if AI agents can act on your site, not just read it

Use tool →
MCP GeneratorLive

Generate your MCP Server Card and a working starter server

Use tool →
AEO Citation MonitorLive

Track if ChatGPT, Claude, Perplexity and Gemini recommend your business

Use tool →
Vertical AI LeaderboardLive

See where any brand ranks in AI-generated responses by category

Use tool →
Fan-Out Query AnalyzerLive

Reveal the hidden queries ChatGPT and Claude fire when researching any topic

Use tool →

New tools ship continuously. Free tier always available.

Browse all tools →