GEO & AEO
Written on 1/9/2026
Updated on 1/9/2026
3min

llms.txt: what it actually does, and what it doesn't

Thibaut Legrand
Thibaut Legrand
Co-founder - Vydera
llms.txt guide Vydera
Table of contents

Can AI engines actually read your site?

Architecture, structured data, llms.txt. Vydera audits and fixes.

Talk to an expert

Key takeaways

  • llms.txt is a markdown file at the root of your site that describes its structure in plain language, for language models to read
  • No major model has officially announced reading it. That is not reason enough to skip it, and not reason enough to expect miracles from it
  • It blocks nothing and forces nobody to cite you. To deny access, that is robots.txt. To list your URLs, that is sitemap.xml
  • Its real value sits elsewhere: writing an llms.txt forces you to state in three sentences what your company does. Plenty of teams discover at that moment that they cannot
  • 7 of 31 industry sites ship a usable one, and 2 French SEO agencies out of 15. The cost is two hours, the risk is zero, the payoff is uncertain but asymmetric

llms.txt was proposed in September 2024 by Jeremy Howard. Two camps formed almost immediately. On one side, the people installing it everywhere and calling it the sitemap of the AI era. On the other, the people shrugging and pointing out that no major model has announced reading it.

Both are wrong, though not where they think.

Vydera runs an llms.txt on vydera.com. This article covers what it does, what it doesn't, how it is built, and what 31 sites in the industry actually do with theirs.

llms.txt in one sentence

It is a markdown file placed at the root of a site, describing in plain language what the site contains and what each group of pages is for.

Nothing more. No markup, no schema, no technical directive. Structured text, readable by a human and by a machine.

The starting idea is simple: a modern web page is a poor reading surface for a model. Between navigation, scripts, consent banners and interface components, the useful content is often under 10% of the HTML served. llms.txt offers a shortcut: here is the map, here is where to go, here is why.

Three things llms.txt does not do

1. It blocks nothing

This is the most common confusion. llms.txt carries no prohibition. A crawler that reads it is under no obligation to honour it, and a crawler that ignores it breaks no rule.

To control access, the file to edit is robots.txt, with the user-agents of the AI crawlers. That is a separate piece of work.

2. It forces nobody to cite you

A well-written llms.txt does not mechanically raise your citation rate. It makes your content easier to understand for a model that has already decided to go and fetch it. That decision is made upstream: brand authority, density of mentions, quality of the sources that talk about you.

We break that mechanism down in our article on how generative AI answers are built.

3. It is not a sitemap

A sitemap lists URLs. An llms.txt explains what they are for. The difference is the description after each link: that is where all the value sits, and that is exactly what most auto-generated llms.txt files leave out.

FileWhat it doesWho honours itBlocks access
robots.txtAllows or denies crawling, user-agent by user-agentEvery serious crawlerYes
sitemap.xmlExhaustive list of URLs with their last-modified dateGoogle, BingNo
llms.txtExplains the site structure and the role of each group of pages, in plain languageNo officially announced adoptionNo

Anatomy of an llms.txt that earns its place

Here is the vydera.com file, zone by zone. Click a zone to see the rule that applies and the most common mistake attached to it.

Anatomy of the vydera.com llms.txt

Click a zone of the file to see the rule that applies to it.

Live file, captured on 27 August 2026 at https://vydera.com/llms.txt

We tested 31 sites in the industry. 7 have a usable llms.txt.

Everyone talks about it, very few actually ship one. On 27 August 2026 we requested /llms.txt on 31 domains: 15 French SEO agencies, 15 SEO and GEO tools, plus vydera.com. Here is the raw result.

7 out of 31 serve a usable file. The other 24 return a 404, a 403, or an HTML page. Among French SEO agencies the score drops to 2 out of 15: noiise.com and uplix.fr. Eskimoz, Primelis, Resoneo, Semji, Keyweo, Digimood, 1ère Position, Slashr, Korleon, Natural Net, Open Linking, seo.fr: nothing.

On the tooling side the contrast is sharper still. Ahrefs, Botify, OnCrawl, Screaming Frog, Babbar, Haloscan and MyPoseo have none. And among the tools that sell AI visibility specifically, Peec AI, Profound, Nightwatch and Authoritas have none either. Only Otterly and Scrunch ship one.

The 7 usable llms.txt files out of 31 domains tested, 27 August 2026
DomainTypeLinksWith descriptionTo markdown
similarweb.comTool11111193
scrunchai.comAEO tool9520
vydera.comFR agency814381
otterly.aiAEO tool49120
semrush.comTool42420
uplix.frFR agency31310
noiise.comFR agency000

The table says two things.

Having a file is not the same as having a good file. noiise.com serves 5,669 bytes across 10 H2 sections and zero markdown links: the file exists and leads nowhere. Scrunch lists 95 links of which only 2 carry a description: the disguised sitemap, the exact mistake described above. Otterly manages 12 out of 49.

The best-kept files come from the tools, not the agencies. Similarweb serves 111 links, every one described, 93 of them pointing at markdown variants. Semrush: 42 links, 42 descriptions. uplix.fr, the only French agency shipping a clean file: 31 out of 31.

And Vydera? 81 links, 43 descriptions. 38 bare links, almost half the file. Better than the industry average, and not clean. The fix is in our backlog, and publishing that beats pretending otherwise.

The markdown variants, at least, are real

One thing is verifiable without access to anyone's logs: what the server returns. Vydera's llms.txt points at 81 /index.md URLs. All of them answer 200, with Content-Type: text/markdown and clean markdown in the body.

Of the 7 usable files in the sample, only 2 do this: Similarweb (93 .md links) and Vydera (81). The other five point at regular HTML pages.

That is where most of the work actually sits. A link to an HTML page weighed down by navigation, scripts and banners forces the model to sift. A link to markdown hands it the content. The file is the entry point, the markdown variants are the product. Doing one without the other is half the job.

Should you write one? The honest answer

Yes, and for a reason that has nothing to do with AI engines.

Writing a decent llms.txt forces three exercises most teams have never done:

  1. Summarise the company in three sentences, with no promotional adjectives, stating what it does, for whom, and with what measurable outcome. That is the blockquote. Plenty of teams discover at that point that they have no clear answer.
  2. Classify pages by role, not by position in the menu. What informs, what sells, what proves, what exists only for legal reasons.
  3. Write a useful description for every page. Not the title copied over: a sentence saying what is there and why anyone would read it.

All three improve the site, the copy and the internal linking, whether AI engines read the file or not.

The maths are simple: two hours of work, zero technical risk, an uncertain payoff on the engine side and a certain one on clarity. That is an asymmetric bet, and it is worth taking.

The five mistakes we see most

  1. The auto-generated file. A plugin that dumps the sitemap into markdown produces a file with no descriptions, therefore no value. llms.txt is written by hand.
  2. The advertising blockquote. "Leader in digital transformation" tells a model nothing. "SEO and GEO agency, B2B clients in France, results measured at 90 days" tells it something.
  3. Relative URLs. The file can be fetched out of context. Always absolute URLs.
  4. The 400-line file. An llms.txt is not exhaustive. It is selective. List everything and you prioritise nothing.
  5. The write-once-and-forget file. An llms.txt pointing at deleted pages is worse than no file at all. Review it at every redesign, like the sitemap.

Where to go next

llms.txt is one brick, not a strategy. It makes sense inside a set: clean structured data, content written to be cited, and an architecture engines can traverse without getting lost.

If you do not know where to start, the question is not "do I have an llms.txt". It is: does a model landing on my site understand within thirty seconds what I do and for whom. If the answer is no, the file will not change that.

For the wider picture of what is at stake on AI engines, our article on GEO / AEO: the full definition sets the ground. And if the topic is site architecture, that is a project of its own.

  • Is llms.txt required to be visible in AI answers?

    No. No major model has officially announced reading the file, and a brand can be widely cited without one. llms.txt makes your site easier to understand for a model that has already decided to read it, it does not trigger that decision.

  • What's the difference between llms.txt and robots.txt?

    robots.txt allows or denies crawling, llms.txt explains the site structure. robots.txt is honoured by every serious crawler and genuinely blocks access. llms.txt carries no prohibition: it is a reading aid, not a directive.

  • Where should the llms.txt file live?

    At the root of the domain, reachable at https://yoursite.com/llms.txt, served as text/plain or text/markdown. Same logical spot as robots.txt and sitemap.xml.

  • Should I also create markdown versions of my pages?

    It is what separates the serious files from the rest. Across the 31 industry domains we tested on 27 August 2026, 7 serve a usable llms.txt, and only 2 point at markdown variants: Similarweb and Vydera. Serving clean markdown saves the model from wading through HTML loaded with navigation and scripts.

  • How long does writing an llms.txt take?

    Around two hours for a mid-sized site, if written by hand. A file auto-generated from the sitemap takes a minute and is worthless, because it has no descriptions.

  • Can an llms.txt hurt classic SEO?

    No. The file is not indexable as a content page, it competes with nothing and changes no Google signal. The only risk is leaving a file online that points at deleted pages, which sends a poor freshness signal.


Thibaut Legrand
Thibaut Legrand
Co-founder - Vydera