AICanary
● Guide

Does llms.txt actually work?

llms.txt is a proposed file that tells AI systems what your site contains. It costs an hour to write and it is widely recommended. It is also widely oversold — no major AI provider has committed to reading it.

Written by PX7 Digital, who publish llms.txt on their own sites and measure what AI answers actually cite.

What llms.txt is

A Markdown summary at a fixed address

A plain file at /llms.txt describing what the site is, what it offers, and which pages matter — written for a language model rather than a browser. The proposal was published in 2024 by Jeremy Howard, and the format is deliberately simple: a heading, a summary line, and lists of links with one-sentence descriptions.

Modelled on robots.txt — but not the same thing

robots.txt is a decades-old standard that crawlers genuinely obey. llms.txt borrows the placement and the plain-text simplicity, but it is a proposal, not a standard, and obedience is the open question rather than the premise.

Descriptive, not restrictive

It does not block anything or grant permission. It says here is what we are and where the good pages are. Access control still belongs in robots.txt and in your provider-specific crawler settings.

What the evidence actually shows

No major provider has confirmed reading it

OpenAI, Anthropic, Google and Perplexity have not committed to using llms.txt. Google has publicly said it is not used for Search. Anyone telling you it drives AI citations is describing a hope, not a documented behaviour.

AI answers cite ordinary sources

In the answer evidence we collect, assistants name conventional sources — app stores, code hosts, established review sites, documentation and the pages themselves. We have not seen an assistant attribute an answer to a site's llms.txt.

It is still cheap and harmless

An hour of work, no ongoing cost, no risk of penalty, and a real chance the convention gains adoption. Writing one is a reasonable bet. Expecting it to change your answers this quarter is not.

What moves AI answers instead

Crawler access to the real content

Check that GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot and Googlebot receive your actual pages and not a JavaScript shell. A site whose content only exists after scripts run is invisible to much of the ecosystem, whatever its llms.txt says.

Facts an assistant can verify

Prices, versions, requirements and claims that agree between your visible copy and your structured data. Assistants that cannot confirm a fact tend to omit it — or invent a plausible substitute.

Presence in the sources they already cite

App stores, package registries, code hosts and independent write-ups. This is slower and harder than adding a file, which is exactly why the file is more popular advice.

Content that answers the question being asked

Assistants summarise pages that address the user's actual question. A page comparing the real options in a category earns citation in a way a product page rarely does.

If you write one anyway

Keep it true and keep it current

A stale llms.txt is worse than none: it hands a confident summary of last year's product to anything that does read it. Update it in the same commit as the change it describes.

Link to pages, not marketing

One line per page, saying what the page answers. The file is a map, not a pitch — and a model reading a pitch has nothing to summarise.

Serve Markdown for the pages too

Content negotiation on Accept: text/markdown, or a plain index.md beside each page, gives any agent the text without the markup. It costs about as little as llms.txt and helps every reader that arrives.

Find out what AI actually says about you.

AICanary checks the technical foundation for free, then runs real customer questions against the major assistants to see whether you appear, who is recommended instead, and which sources shape the answer.

Run the free report