Guides

What is llms.txt, and should you publish one?

llms.txt is a map for AI agents, not a ranking signal. What the v2 proposal asks for, what the evidence says about who reads it, and how to make yours findable.

By The Laarpi teamUpdated 7 min read

llms.txt is a Markdown file at the root of a website, at /llms.txt, that tells AI agents what the site is and which pages to read, ideally as clean Markdown versions of those pages. Jeremy Howard proposed it in September 2024, and the current version, v2, is dated 10 August 2026. It is a map for agents, not a ranking signal: Google Search says it doesn't use the file, while agents do use it when something points them to it. That is why the most useful thing in v2 is a way to make the file findable. It is one part of an AI-ready website, next to readable server HTML, structured data and forms that agents can use.

The short answer

Three findings decide whether to publish one, and they are all true at once:

  1. Google Search doesn't use it. Google's AI optimization guide says: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." A note Google added on 15 June 2026 says these files won't help or hurt your visibility or rankings, and that it's fine to keep them for other systems that use them.
  2. Most files are never fetched. Ahrefs looked at 137,210 domains with traffic in May 2026. 28% of them publish an llms.txt, and 97% of those files received no requests at all that month. AI bots made zero requests for llms.txt files that didn't exist: "They never go looking."
  3. When a page links to it, agents waste far fewer requests. Mintlify, a docs platform that sells llms.txt support, ran 2,400 tasks with Claude Code and Codex on 20 docs sites. The average number of 404s per task fell from 2.23 with HTML pages, to 1.42 with plain Markdown, to 0.11 when the Markdown linked to an llms.txt.

So publish one if agents might want to use your site, keep it short and true, and link to it from the pages agents already fetch.

What the file looks like in v2

The proposal fixes the order of the parts:

  1. An H1 with the name of the site. This is the only required part.
  2. A blockquote with a short summary holding the facts an agent needs to understand the rest.
  3. Any Markdown except headings: paragraphs or lists that explain the site and how to read the files.
  4. H2 sections of file lists. Each item is a Markdown link [name](url), optionally followed by a colon and a note.

A section called ## Optional is a convention for links an agent can skip when it needs a shorter context. In v1 it had a mechanical meaning for context-building tools. v2 drops that, so it is now only a label.

Here is a complete example for a small studio site. Copy it, then replace every line with facts that are true for you:

# Northlight Studio

> Northlight is a two-person brand and web studio in Lisbon. We design
> identities and websites for food and hospitality businesses.
> Prices are in euros and exclude VAT.

Every page below has a Markdown version at the same address, ending in .md.
For availability and quotes, use the contact page rather than email.

## Services

- [Brand identity](https://example.com/services/brand-identity.md): logo, type, palette and guidelines, with the starting price
- [Websites](https://example.com/services/websites.md): design and build, what the client can edit afterwards

## Work

- [Case studies](https://example.com/work.md): each project's brief, result and client

## Actions

- [Book an intro call](https://example.com/contact.md): the form's fields, what happens next and the reply time

## Optional

- [Journal](https://example.com/journal.md): essays on type and menus

Write the blockquote for a reader who has never heard of you and has one question. Prices, places, opening hours and what you don't do are worth more than a slogan.

The Markdown twins

The file lists work best when they point at clean Markdown rather than full HTML pages with navigation, cookie banners and scripts. v2 says pages that agents might need "provide a clean markdown version of those pages at the same URL as the original page", in either of two forms:

PageTwin, v1 form (still valid)Twin, v2 form (new)
/services/websites.html/services/websites.html.md/services/websites.md
/about (no extension)/about.md/about.md

A twin carries the page's content: headings, text, prices, lists and links. It leaves out the chrome. It must say the same thing as the page. A twin that still shows last year's price is worse than no twin, because an agent will quote it.

There is another route if your site sits behind Cloudflare. Its Markdown for Agents feature converts HTML to Markdown at the edge when a client sends Accept: text/markdown. It is off by default, available on Pro, Business and Enterprise plans, works only on HTML responses up to 2 MB, and adds an x-markdown-tokens header with an estimated token count. That gives agents Markdown without you maintaining twins, though the conversion is automatic rather than written.

Ahrefs' study explains why a file at the root often goes unread: "AI tools fetch llms.txt when a link, an index, or a user instruction tells them it exists." v2 adds two standard link relations so that every page can point to its own Markdown version and to the llms.txt that covers it.

In the <head> of each page:

<link rel="alternate" type="text/markdown" href="/services/websites.md">
<link rel="describedby" href="/llms.txt">

Or as an HTTP header, which also reaches clients that never parse HTML:

Link: </services/websites.md>; rel="alternate"; type="text/markdown", </llms.txt>; rel="describedby"

Two ways to send the describedby header on every response:

# nginx, inside the server block
add_header Link '</llms.txt>; rel="describedby"' always;
# Netlify: a file named _headers in the publish directory
/*
  Link: </llms.txt>; rel="describedby"

The alternate link differs per page, so it is easier to write in the HTML template than in server config. Mintlify's benchmark also tested putting the whole llms.txt inside each page. It reduced 404s about as much as the link did but cost more tokens, so the link is the better choice.

Path coverage: one file or several

v2 defines what an llms.txt covers: "A file covers the URLs under its path, and where more than one file applies, agents should use the most specific one." A file at /docs/llms.txt covers /docs/ and everything below it, and an agent reading /docs/api/auth should prefer it over /llms.txt.

This matters in two cases. Large sites can give each section its own short file instead of one file nobody can scan. And sites that control only a path, such as a GitHub Pages project at user.github.io/project/, can publish a valid llms.txt for their part of the domain.

OpenAI's developer site is a real example of the first case, as we found it on 6 October 2026. Its root file at developers.openai.com/llms.txt is a short index whose links point to one llms.txt per section, such as /api/llms.txt and /codex/llms.txt. Each docs page also carries a line for agents saying where the index is and that a Markdown version of the page is available by adding .md to its address. That line is plain text rather than the v2 link tags, but it does the same job: it tells the agent that the map exists.

What the evidence says, and what it doesn't

SourceWhat was measuredWhat it foundRead it with
Google AI optimization guide (updated 10 July 2026)Google Search's own systemsGoogle Search doesn't use llms.txtIt speaks for Google Search only, not for agents
Chrome's Lighthouse, Agentic Browsing category (updated 5 May 2026)Deterministic checks for agent useChecks whether llms.txt is presentExperimental, needs Chrome 150 or later; a Google product that does look for the file
Ahrefs (15 June 2026)137,210 domains, May 2026 traffic28% publish a file; 97% of those got no requests; no AI bot asked for a missing fileIts customers "skew more technical and SEO-aware than the web at large", and a fetch isn't a read
SE Ranking (7 November 2025)Nearly 300,000 domains and their AI citations10.13% had a file; no measurable link between having one and being citedIts own sample; don't set its 10% against Ahrefs' 28%
Mintlify (17 July 2026)2,400 agent tasks on 20 docs sites404s per task: 2.23 HTML, 1.42 Markdown, 0.11 with a linked llms.txtA vendor's benchmark, on docs sites, with coding agents

None of this shows that an llms.txt gets you cited more often. What it shows is narrower and more useful: agents that are pointed to the file use it, and they make fewer wrong turns when they do.

Check whether anything reads yours

If you can read your access logs, this counts requests for the file by user agent (combined log format, where the user agent is the sixth field when split on quotes):

grep '"GET /llms.txt' access.log | awk -F'"' '{print $6}' | sort | uniq -c | sort -rn

If a CDN serves the file from its cache, your origin logs miss most of the requests, so read the CDN's logs instead. And keep Ahrefs' caveat in mind: a request proves that something fetched the file, not that a model read it.

Mistakes worth avoiding

  • A root file that nothing links to. Add the describedby link, or the file waits for an agent that already knows it exists.
  • Twins that drift. Generate the twins from the same source as the pages, or check them every time a price or page changes.
  • Marketing copy in the blockquote. An agent needs facts it can repeat, not adjectives.
  • Listing every URL. That is the sitemap's job. llms.txt is a curated reading list.
  • Blocking the file. A robots.txt rule or a bot challenge on /llms.txt stops compliant agents from reading it at all. How AI agents read websites explains which crawlers, fetchers and browser agents ask for files like this one.

On a Laarpi site

Laarpi Surface writes this layer on every publish, from the same source as the pages: an llms.txt in the v2 shape, a .md twin of every public page, both discovery links in the initial HTML and the same links as HTTP Link headers. Because the pages and the twins come out of one build, a changed price shows up in both at once. What an AI-ready Laarpi site includes lists the rest of the agent door, and structured data for AI search covers the JSON-LD that goes with it.

Sources

Checked 6 October 2026. Standards and vendor pages change; the linked pages are the authority.

  1. llms.txt proposal, v2 (Jeremy Howard, updated 10 August 2026)
  2. llms.txt: changes in v2
  3. Google Search Central: AI optimization guide (updated 10 July 2026)
  4. Google Search Central: AI features and your website (updated 10 December 2025)
  5. Google Search Central: documentation updates (llms.txt note, 15 June 2026)
  6. Chrome for Developers: Lighthouse Agentic Browsing (updated 5 May 2026)
  7. Ahrefs: llms.txt study of 137,210 domains (Louise Linehan, 15 June 2026)
  8. Mintlify: llms.txt agent benchmark (vendor study, 17 July 2026)
  9. SE Ranking: llms.txt study (Yulia Deda, 7 November 2025)
  10. Cloudflare: Markdown for Agents (updated 13 July 2026)
Questions

Fair questions

Does llms.txt help my Google ranking?

No. Google's AI optimization guide says Google Search doesn't use llms.txt or other AI text files, and a June 2026 note in Google's documentation updates says they won't help or hurt your visibility or rankings. Publish one for the agents that use it, not for Google Search.

Do ChatGPT, Claude and Perplexity read llms.txt?

Their crawlers aren't documented as reading it: OpenAI's and Anthropic's crawler pages describe each bot without mentioning the file. What the logs show is that agents fetch it when something points them to it. In Ahrefs' May 2026 data, AI bots never requested an llms.txt that didn't exist, and in Mintlify's benchmark the coding agents made far fewer wrong requests when the page linked to one.

Is llms.txt an official standard?

No. It is a proposal, first published by Jeremy Howard in September 2024 and revised as v2 on 10 August 2026. No standards body like the IETF or the W3C has adopted it. Its v2 discovery links reuse link relations that already exist, which is why they work with ordinary HTML and HTTP.

What is the difference between llms.txt, robots.txt and sitemap.xml?

robots.txt says which crawlers may fetch which paths. sitemap.xml lists every URL you want indexed, for search engines. llms.txt is a short, curated Markdown page that tells an agent what the site is and which few pages to read for which question, ideally as clean Markdown. They do different jobs, and a site can have all three.

What about llms-full.txt?

It isn't defined in the llms.txt proposal. It is a convention some documentation platforms use for a single file holding the full text of every page. It can help on docs sites that agents read end to end, but it isn't required, and a marketing site rarely needs it.

How do I know whether anything reads my llms.txt?

Look for requests to /llms.txt in your server or CDN logs and group them by user agent. Ahrefs' own caveat applies: a fetch shows that something asked for the file, not that a model read it or used it in an answer.

Related

Start building

Describe the site in a sentence. It asks what matters, then designs and builds it from scratch.

One sentence is enough.