llms.txt is a markdown file, served at /llms.txt, that gives AI agents a short, curated map of a website and links to clean versions of its pages. Think of it as a table of contents written for a model instead of a search engine.
Key takeaways
- llms.txt is a proposal, not a formal standard. It was published by Jeremy Howard on 3 September 2024 and last modified on 10 August 2026 (llmstxt.org, September 2026).
- The file is markdown with a fixed order: an H1, a one-line summary, optional notes, then H2 lists of links.
- It is read on demand by an agent that needs help with your product. It is not a ranking signal.
- If you sell an API, it is a cheap file to publish. LeadOcean serves one at
api.leadocean.io/llms.txt.
What it is
llms.txt is a markdown file at the root of a site (or at any subpath) that summarizes the site for language models and lists links to the pages worth reading.
Web pages are built for people. They carry navigation, scripts and banners around the useful text. An agent that fetches a docs page has to strip all of that, and it has limited context to spend.
The proposal fixes this with one small file. It stays short enough to fit in context. The detail sits behind the links, and the agent fetches only what it needs.
The spec also suggests a clean markdown copy of key pages, at the same URL with .md added. The links in llms.txt should point to those copies when they exist.
A second file is common in practice: llms-full.txt. It holds the full reference in one document, for agents that should read everything. Both names are conventions you will see on API docs sites.
How it works
An agent reads llms.txt in four steps. You write the file once and the agent does the rest.
- The agent needs context. A user asks a coding agent to call your API, or a chat assistant to answer a question about your product.
- It fetches
/llms.txt. The file is small, so the whole thing lands in context. - It picks links. The one-line descriptions tell it which page answers the question.
- It fetches only those pages. Ideally these are the
.mdversions, with no page chrome to strip.
The file follows a precise structure, in this order (llmstxt.org, September 2026):
- An H1 with the name of the project or site. This is the only required part.
- A blockquote with a short summary.
- Optional paragraphs or lists with notes on how to read the files.
- H2 sections that hold lists of links, each as
[name](url)with an optional: note.
An H2 called Optional is, by convention, the section an agent can skip when it needs less context.
Here is a worked example for a made-up company. Acme sells an invoicing API.
```markdown # Acme
> Acme is an invoicing API. REST, JSON, API key in the Authorization header. Base URL: https://api.acme.com
Pricing is on the pricing page. The rate limit is 50 requests per second.
## Docs
- Quickstart: create and send your first invoice
- API reference: every endpoint, parameter and error code
## Optional
- Changelog: releases since 2024
```
That file is about 60 words. An agent can read it in one pass and know which two pages to fetch next.
llms.txt vs robots.txt and sitemap.xml
llms.txt guides an agent to content. robots.txt sets access rules and sitemap.xml lists pages for search engines. The three files do different jobs and sit side by side.
| llms.txt | robots.txt | sitemap.xml | |
|---|---|---|---|
| Format | Markdown | Plain text rules | XML |
| Written for | Language models and agents | Crawlers | Search engines |
| Job | Curated overview with links to clean content | States what automated tools may access | Lists indexable pages |
| Used | On demand, when an agent needs an answer | By crawlers before they fetch | By crawlers to find pages |
| Size | Small enough to fit in context | Short | Can list every page |
| Required | No, a proposal | No, a convention crawlers honor | No |
The spec is explicit about the difference with the other two. robots.txt tells tools what access is acceptable. llms.txt is used on demand, when an agent needs information while helping a user. A sitemap lists every page, which usually adds up to more than a context window holds.
The spec adds that llms.txt is expected to help inference more than training, and that this is how it has been used.
When it matters
The file pays off when an agent has to get something right from your docs.
You publish an API
This is the strongest case. A coding agent that reads a clean API reference writes better calls. It uses the right paths, parameter names and auth header. Without it, the agent guesses from older training data.
You have docs that change
Models learn from a snapshot. Your docs move on. A pointer to current markdown pages gives the agent a way to read today's version instead of a remembered one.
You want an agent to explain your product
A sales assistant asked about your pricing and limits will read what it can find. If you list the pricing and policy pages, the agent has a short route to the right ones. This only helps when the linked pages are accurate and current.
You hope it will lift your search ranking
Do not expect that. The proposal describes agents using the file at the moment of a task. It does not describe ranking. Treat it as a hint to agents, and do not promise any particular crawler will fetch it.
How LeadOcean handles it
LeadOcean publishes llms.txt and llms-full.txt for its API host. Both are public and need no key. There is no llms.txt filter in the data, so the field that maps to this idea is the API's own machine-readable surface.
The root file at api.leadocean.io/llms.txt is about 2.9 KB (checked 2 October 2026). It follows the format above: an H1 (LeadOcean), a one-line summary, then sections for Docs, Connect, Endpoints and Price. It links to three deeper resources:
llms-full.txt: every endpoint, parameter, enum value, response schema, error and the MCP tool catalog. It is about 88 KB on the same date.openapi.json: the same surface for a code generator./mcp/tools: the MCP tool catalog as JSON. The server has 12 tools.
Note where the file lives. It covers the API host. The marketing site at leadocean.io does not serve one (checked 2 October 2026).
You can read all three with no key. Fetch them like any file:
curl -s https://api.leadocean.io/llms.txt
curl -s https://api.leadocean.io/llms-full.txt | head -40
curl -s https://api.leadocean.io/mcp/toolsAn agent that reads these knows the canonical endpoint paths, such as POST /v1/people/enrich and POST /v1/people/search. It also learns that sizing is free: count=true with limit=1 returns a total and spends no records. That matters because search costs 1 record per person returned.
The other route is the hosted MCP server at https://api.leadocean.io/mcp. It signs in with OAuth 2.1 and needs no pasted key. Use MCP when the agent should call the data. Use llms.txt when the agent should read how the API works. The Python scraper cheat-sheet covers the related job of turning HTML into markdown for a model.
Free is 1,000 records, one-off, no card. Pro is $499 a month, flat. See pricing.
FAQ
Is llms.txt an official standard?
No. It is a community proposal published at llmstxt.org. Thousands of sites publish one, and the proposal says that AI labs publish them for their own developer docs (llmstxt.org, September 2026). There is no standards body behind it, so support varies by tool.
Where do I put the file?
At the site root as /llms.txt, or at any subpath such as /docs/llms.txt. A file covers the URLs under its path. Where more than one applies, agents should use the most specific one.
What is the difference between llms.txt and llms-full.txt?
llms.txt is the short map with links. llms-full.txt is the whole reference in a single file. The short file fits any context window. The full file suits an agent that should read everything before it writes code.
Does llms.txt replace robots.txt?
No. robots.txt states what automated tools may access. llms.txt points agents at the content you want them to read. Keep both.
How do I check that mine works?
The proposal gives one test. Ask an agent questions about your product, and give it only your llms.txt as a starting point. If it finds the right pages, the descriptions are doing their job.
Give your agent B2B data it can read and query. Start free.
Free to start. No credit card. 1,000 records to spend whenever you like.
Get your free API key →