LIVE MONITORINGChatGPT · Claude · Gemini · Grok · DeepSeek · Mistral · Perplexity · AI Overviews (Google) · AI Mode (Google) · Copilot (Microsoft) · Meta AI11 ENGINESv.2026.04 / build 47GENERATIVE ENGINE OPTIMIZATIONLISBON · PORTUGAL
Technical standards8 min read

What llms.txt is and how to create yours

A practical specification of llms.txt, the markdown file at the root of the site that AI crawlers read before the HTML. Structure, rules, a real example and the mistakes to avoid.

PTLer em português →

Photo by Markus Spiske on Unsplash

llms.txt is a markdown file at the root of the site that describes the content in a structured, language-model-readable way. An emerging convention in 2024 to 2026, adopted by Anthropic, Mistral and several technical companies. It coexists with robots.txt and sitemap.xml, it does not replace them. This piece shows the minimum structure, the expanded llms-full.txt version, and the typical mistakes.

Update · June 2026

Google published in June 2026 that it is not necessary to create an llms.txt in order to appear in AI answers in its search. That does not contradict this article: we always described it as low-cost hygiene, not as a citation lever. It still makes sense to publish it (documentation, engines other than Google, zero cost). The full context of what changed is in Google made GEO official.

Key takeaways

What llms.txt is

llms.txt is a public markdown file, hosted at the root of the domain (https://example.com/llms.txt), that describes the structure and purpose of the site in language optimised for models. The specification was proposed in 2024 and defined three fixed blocks: name, a one-line description, and sections of curated links grouped by topic.

The intent is to solve a concrete problem. When an AI crawler arrives at your site, it has to infer the structure through HTML, headers, menus, CSS classes. That process is fragile. With llms.txt, it receives a declarative summary, in a format the models extract with near-total fidelity.

Why it matters for GEO

Three practical reasons. First, fidelity: structured markdown has less noise than HTML rendered with JavaScript. Second, curation: the choice of what goes in is yours, instead of leaving the crawler to guess. Third, single fetch: the llms-full.txt version lets the model pick up the whole context in one call, instead of making 10 or 20 subsequent requests.

In GEO terms, this shows up in what the engines cite. The cleaner the representation of your site, the more likely the citation is to be correct: right URL, right name, right description.

Minimum structure

The official specification is simple. Title as h1, description as blockquote, and h2 sections with lists of links. A real example, adapted from the llms.txt of destaque.ai:

# destaque.ai

> Empresa de tecnologia que mede a visibilidade das marcas nas respostas de IA.
> Opera a presença de marcas digitais e locais em ChatGPT, Claude, Gemini e Perplexity.

## Serviço
- [Como trabalhamos](https://www.destaque.ai/servico)
- [Glossário GEO](https://www.destaque.ai/glossario)

## Empresa
- [Sobre](https://www.destaque.ai/sobre)
- [Contacto](https://www.destaque.ai/contacto)

## Blog
- [Posts](https://www.destaque.ai/blog)

## Recursos para IA
- [llms-full.txt](https://www.destaque.ai/llms-full.txt)
- [ai.txt](https://www.destaque.ai/ai.txt)

Note three choices. The links are absolute, not relative. The sections are grouped by intent (service, company, blog), not by page type. And there is an explicit resources-for-AI section pointing at llms-full.txt, which makes the crawler's life easier.

llms.txt vs llms-full.txt

llms.txt is the index. llms-full.txt is the index plus the expanded body of the key pages, concatenated in markdown.

In practice: the crawler starts with llms.txt to map the terrain; if it wants depth, it goes to llms-full.txt and picks up everything at once. The cost of serving llms-full.txt is low (a static route) and the saving in crawler requests is considerable.

Rule of thumb: if you already have 5 well-written key pages, it is worth generating the expanded version. If you are still building content, focus on llms.txt first.

The index and the expanded version

llms.txtllms-full.txt
What it isThe indexThe index plus the expanded body of the key pages, concatenated in markdown
What the crawler doesStarts here to map the terrainComes here for depth and picks up everything at once
Size30 to 100 lines is enough for most companiesUp to 50 to 100 thousand tokens, depending on the size of the site
WhenFirst, if you are still building contentOnce you have 5 well-written key pages

llms.txt vs ai.txt vs robots.txt

Four files, four distinct responsibilities:

Four files, four responsibilities

robots.txtsitemap.xmlllms.txtai.txt
GovernsAccessIndexingUnderstandingThe AI usage policy
What it doesIt tells the crawlers (including GPTBot, ClaudeBot and the rest) what they may or may not crawl.It lists the site's URLs so crawlers discover pages.It explains in markdown what each block of the site is.It specifies which AI crawlers may index and on what terms. Analogous to robots.txt but focused on use by generative models.

All four should coexist. Each solves a different problem in the AI visibility work.

How to create yours, in five steps

The process is shorter than it looks. Typically an afternoon's work.

Five steps

01 · Identify the pages that matter

Typically: home, service, about, contact, blog index. For most B2B SaaS sites, that is 5 to 10 pages.

02 · Write a descriptive sentence for each

Short, specific. "How we work: the four phases of the method" is better than "Service page".

03 · Group by intent, not by page type

Service, company, resources, blog. Not static pages and dynamic ones.

04 · Generate the file as a static route or route handler

In Next.js, a route handler at app/llms.txt/route.ts does it. In other stacks, a static file in public/.

05 · Reference it in robots.txt

A simple line: Sitemap: https://example.com/llms.txt. It is not a formal sitemap, but it gives extra discovery. Repeat for llms-full.txt.

Typical mistakes

We see the same four again and again:

How to validate

There is no official validator yet. The manual process is simple:

Validating, by hand

Frequently asked questions

Does llms.txt replace sitemap.xml or robots.txt?

No. They are three files with different purposes. robots.txt says who may enter; sitemap.xml says where the pages are for indexing; llms.txt says, in LLM-readable markdown, what each page is and how it is organised. They coexist.

Do AI crawlers actually read llms.txt?

Adoption is growing but still partial. Anthropic, Mistral and some large technical companies have adopted it. OpenAI and Google have not confirmed official support. Even so, the effort is low and the upside scenario is large, so it is worth putting in place.

Do I have to create llms-full.txt as well?

It is recommended, yes. llms.txt works as an index; llms-full.txt carries the expanded body in markdown, the main text of the key pages concatenated. A single fetch for the crawler to resolve everything in one call.

How big should llms.txt be?

Short. llms.txt is a curated index, not an exhaustive sitemap. For most companies, 30 to 100 lines is enough. llms-full.txt can go up to 50 to 100 thousand tokens depending on the size of the site.

Sources