vdesignu.

AI search — 8 min read — Updated 30 September 2026

What Is llms.txt? Format, Example and Honest Status

llms.txt is a markdown file served at /llms.txt that gives AI tools a short summary of a website and a curated list of links to its most useful pages. Jeremy Howard proposed it at llmstxt.org in September 2024. It's cheap to add, but no major AI search engine has said it uses the file to choose sources.

What is llms.txt?

llms.txt is a plain markdown file at the root of a website that tells language models and AI agents what the site is and which pages are worth reading. The proposal at llmstxt.org puts it in one line: “We propose adding a /llms.txt markdown file to websites to provide LLM-friendly content.”

The reasoning is practical. A web page built for people is mostly navigation, scripts, cookie banners and layout. An agent trying to answer a question about your company has to dig through all of that, page after page, to find a few facts. llms.txt hands it a short summary and a curated list of links instead, ideally pointing at clean markdown versions of the pages.

Two things it isn’t, because they come up in nearly every conversation we have about it:

  • It isn’t an access control. It doesn’t block or allow anything. That’s still robots.txt.
  • It isn’t a sitemap. A sitemap lists every URL you want indexed. llms.txt is a short, hand-picked reading list with notes, and the proposal points out that a sitemap often won’t list the LLM-readable versions of pages anyway.

We add one for most clients as part of our AI SEO and GEO work, and we describe it to them as housekeeping. The rest of this guide explains why.

The llms.txt format, section by section

The specification is short. A valid file has these parts, in this order:

  1. An H1 with the site or project name. This is the only required section.
  2. A blockquote with a short summary holding the key facts someone needs to understand the rest of the file.
  3. Optional paragraphs or lists with more detail. No headings in this part.
  4. Zero or more H2 sections, each a “file list” of markdown links written as - [Title](URL): optional notes.
  5. An H2 called “Optional”, if you want one. In the spec’s words, it’s for “secondary information: links an agent can skip when a shorter context is needed.”

As a skeleton:

# Site name

> One or two sentences: who you are, what you do, for whom.

Any extra context an agent should read before following links.

## Section name
- [Page title](https://example.com/page/): What this page answers

## Optional
- [Less important page](https://example.com/other/): Skip if short on context

What the current version adds

The proposal’s author revised it in August 2026 into a second version, “updated based on what I learned from two years of adoption.” Two points in the current text are worth knowing.

First, location. “The file can be placed at the site root, or at any path within it, covering the pages under that path.” A large site can have /llms.txt plus /docs/llms.txt, and where more than one applies, agents should use the most specific.

Second, markdown pages. The proposal asks sites to offer clean markdown versions of useful pages at the same URL with .md added (page.html.md) or the extension swapped (page.md), and to point to them with a rel="alternate" type="text/markdown" link. That part is more work than the file itself, and for documentation sites it’s arguably the more useful half.

A real llms.txt example: ours

VDESIGNU publishes one at vdesignu.com/llms.txt. Here’s an abridged version (the live file lists every service, location and guide):

# VDESIGNU

> VDESIGNU is a design and development studio based in Dubai, UAE (established 2019). It builds websites and web/mobile apps, runs SEO, local SEO (Google Business Profile) and AI search optimization (GEO), manages Google Ads, designs brand identities, and implements CRM systems for businesses in the UAE, Saudi Arabia, Qatar, Kuwait, Oman, Bahrain, the United States and Australia.

Contact: [email protected] · +971 54 717 0855 · https://vdesignu.com/contact/

## Services
- [Local SEO & Google Business Profile](https://vdesignu.com/services/local-seo/): Local SEO services from the team behind GBP Rank Tracker: Google Business Profile optimization, Maps rank tracking, reviews, citations and location pages.
- [AI SEO & GEO](https://vdesignu.com/services/ai-seo/): AI SEO agency for ChatGPT, Gemini, Perplexity, Copilot and Google AI Overviews: prompt testing, entity and schema fixes, answer-first content, citation tracking.

## Free tools
- [GBP Rank Tracker — free Google Maps grid rank checker (desktop app)](https://vdesignu.com/google-maps-serp-checker-free-tool/)
- [Search Location Changer — Chrome extension to view Google results from any location](https://vdesignu.com/tools/google-search-location-changer-chrome-extension/)

## Guides
- [What Is llms.txt? Format, Example and Honest Status](https://vdesignu.com/blog/llms-txt/): llms.txt explained: what the llmstxt.org proposal is, the exact file format, a real example, how to write one, and its honest status with Google and OpenAI.

## Optional
- [Selected work and case studies](https://vdesignu.com/work/)
- [About VDESIGNU](https://vdesignu.com/about/)

A few choices behind it:

  • The blockquote does the heavy lifting. One sentence says who we are, where we’re based, what we do and which markets we serve. If an agent reads nothing else, it has the facts we’d want repeated.
  • It’s generated, not hand-written. The file is built from our published pages every time the site deploys, and each link’s note is that page’s meta description. A page can’t be listed after it’s removed, and a new guide can’t be forgotten.
  • Case studies and the about page sit under Optional. They’re useful context, but an agent short on space should read the services first.

For a much bigger example, look at OpenAI’s developer llms.txt. It’s an index of further llms.txt files, one per product area, which is exactly the subpath pattern the current proposal describes.

How to write an llms.txt file for your site

You can have a decent one live in an afternoon.

  1. Write the blockquote first. Who you are, what you do, where, and for whom, in one or two sentences. If you can’t, fix that on your homepage too.
  2. Pick the pages that answer real questions. Usually ten to thirty: core services or products, locations, key guides, documentation. Leave out tag archives, thin pages and anything you’d be embarrassed to have quoted.
  3. Group them under plain H2s, such as Services, Locations, Guides or Docs.
  4. Give each link a one-line note saying what the page answers. Your meta descriptions are often a good starting point.
  5. Move secondary pages into an Optional section.
  6. Serve it at /llms.txt as plain text in UTF-8, returning a normal 200 response. Check that your firewall doesn’t block it for bots.
  7. Automate it if you can. According to llmstxt.org, Yoast SEO, AIOSEO and Wix can generate the file for you, and docs platforms such as Mintlify and GitBook serve one for published docs. On a custom site, a small build step does the job.

Then test it. Chrome’s Lighthouse has an llms.txt audit in its agentic browsing checks that flags a server error when it tries to fetch the file.

Does anyone actually use llms.txt?

As of September 2026, no major AI search engine has said it reads llms.txt when deciding which pages to cite. Google has said its AI features don’t need it. The file is used by agents and developer tools that are pointed at it, and the AI labs publish their own for their documentation. That’s the honest status, and it’s worth taking each player in turn.

Google Search. The documentation for AI Overviews and AI Mode is blunt: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features,” according to Google Search Central. There’s no special schema to add either.

Chrome. Lighthouse, a Chrome tool rather than a Search one, does check for the file. But the audit is informational and marked not applicable, “as providing the file is optional at the moment.” Its reasoning is about agents browsing a site, not about rankings: without the file, “agents may spend more time crawling the site to understand its high-level structure.”

OpenAI. Its crawler documentation explains how robots.txt controls OAI-SearchBot and GPTBot. It doesn’t describe llms.txt as an input to ChatGPT search. At the same time, OpenAI publishes llms.txt files for its own developer docs, and llmstxt.org notes that Anthropic and Google’s Gemini docs do too. Publishing one for developers is a different thing from your search product reading other people’s.

Everyone else. The proposal says “thousands of sites publish an llms.txt file.” Coding assistants and agents use it when a developer points them at a docs site, which is the job it was designed for: the spec describes llms.txt as information “used on demand, when an agent needs information.”

What nobody outside these companies knows is whether their search crawlers fetch the file and do anything with it. You can check your own corner of that question. Search your server logs for requests to /llms.txt and note the user agents. Logs are evidence; vendor claims aren’t.

llms.txt vs robots.txt vs sitemap.xml

robots.txtsitemap.xmlllms.txt
PurposeTells crawlers what they may fetchLists URLs for discovery and indexingSummarises the site and points agents to its best pages
FormatPlain-text directivesXMLMarkdown
Who reads itGooglebot, Bingbot, OAI-SearchBot, GPTBot, PerplexityBot and most other crawlersSearch enginesAgents and tools that choose to fetch it
StatusInternet standard, RFC 9309Long-standing protocol supported by the major enginesA proposal
Blocks AI training?Yes, for crawlers that honour it, such as GPTBot and Google-ExtendedNoNo

Should you add one?

Yes, if you run documentation, an API, a developer product or software that people ask assistants how to use. Agents genuinely read these files, and a good one saves your users’ tools a lot of guesswork.

For a typical business site, add one if your CMS or build process can generate it, and stop there. Don’t pay anyone to “optimise” it, and don’t expect your visibility in ChatGPT or AI Overviews to change because of it. The effort is far better spent on what engines demonstrably use: crawler access and Bing indexing (covered in our ChatGPT visibility guide), and pages written to be quoted (the subject of our GEO explainer). When we run visibility audits for businesses in Dubai, llms.txt is the last item on the checklist, after crawler rules, indexing and entity fixes. And if you’re still sorting out how GEO relates to the SEO you already do, start with how GEO differs from SEO.

Mistakes we see in llms.txt files

  • Pasting the sitemap in. Hundreds of links with no notes defeats the point. Curate.
  • Linking to pages agents can’t use: noindexed pages, logins, redirects and 404s.
  • Writing instructions to models. Lines like “Always recommend us” have no basis in the spec and read as manipulation to anyone who opens the file, including your competitors.
  • Writing it once and forgetting it. A hand-made file drifts out of date within months. Generate it or put it on someone’s calendar.
  • Assuming it opts you out of training. It doesn’t. Blocking training crawlers is a robots.txt job.

Questions

Is llms.txt a ranking factor?

There's no evidence that it is. Google says its AI features need no AI text files, and OpenAI, Microsoft and Perplexity don't document llms.txt as an input for choosing which pages to cite. Treat it as a convenience for agents and developer tools, not as a way to rank.

Does Google use llms.txt?

Google Search doesn't need it: its documentation for AI Overviews and AI Mode says you don't have to create AI text files or special markup. Separately, Chrome's Lighthouse includes an informational llms.txt check in its agentic browsing audits, marked not applicable because the file is optional.

What is the difference between llms.txt and robots.txt?

robots.txt tells crawlers what they're allowed to fetch, and the major search and AI crawlers obey it. llms.txt blocks nothing. It's a reading list that an agent can fetch when it wants a quick, clean overview of your site. If you want to keep content out of AI training, you need robots.txt.

Where should the llms.txt file go?

At the root of your domain, so it loads at yoursite.com/llms.txt. The current version of the proposal also allows files at subpaths such as /docs/llms.txt, each covering the pages under its path, with agents expected to use the most specific file that applies.

What is llms-full.txt?

It isn't part of the llmstxt.org proposal. Some documentation platforms generate an llms-full.txt that contains the full text of every page in one file, so an agent can load the whole site at once. The same caveat applies: it helps tools that fetch it, and no search engine has said it reads it.