Skip to content
Issue docs

llms.txt does not follow the llms.txt format

Standardllms_txt_malformedIssue 23

What is this issue?

Your site publishes an llms.txt file, but the file does not follow the format described at llmstxt.org.

llms.txt is a Markdown file at the root of your site — yoursite.com/llms.txt — that hands AI assistants a curated index of your most important pages. It is the AI equivalent of a well-organised table of contents: rather than making a model guess what matters on a page full of navigation, banners and footers, you tell it directly.

The format is small and specific:

## Acme Shoes

> We have sold running shoes online since 2011.

### Products
- [Running shoes](https://acme.com/running): Our main range
- [Trail shoes](https://acme.com/trail): For off-road

### Company
- [About us](https://acme.com/about): Who we are

This finding is raised when the file exists but is missing something the format requires:

  • No H1 title — the file does not open with a # Site name heading.
  • No links — the file contains no page URLs at all, so it indexes nothing.
  • Too short — the body is under 50 characters, which cannot be a usable index.

What does not fail the check

Several things are reported as recommendations rather than as violations. They appear in the finding's detail so you can act on them, but none of them raises this issue on its own:

  • A missing > summary blockquote.
  • Links not grouped under ## Section headings.
  • Links written as plain URLs — - Home: https://site.com/ instead of - [Home](https://site.com/). The Markdown form is what llmstxt.org specifies, and it is genuinely better because it gives each entry a name an assistant can choose by. But a file listing plain URLs is a working index, and failing it would be failing punctuation.
  • Sections marked with # instead of ##. Accepted when the file uses no ## anywhere, since the structure is clearly intended.

This distinction was drawn after measuring the check against real files. Two of eleven surveyed — including a 206-entry documentation index — used plain URLs throughout. Calling those files "malformed" would have been a confidently wrong statement about indexes that work.

What this is not

This is not about the file being missing — that is issue #22, llms_txt_file. This issue only ever fires when the file is present and readable, and it describes the state of its contents.

Why it matters

A malformed llms.txt is worse than no llms.txt at all.

When the file is absent, an AI assistant falls back to reading your pages normally — imperfect, but it works. When the file is present but unparseable, the assistant fetches it, fails to extract anything useful, and moves on. You have spent the effort of publishing a file and gained nothing from it, while believing the job is done.

Each of the three required elements does specific work:

The # Title heading tells the model whose index it is reading. Without it there is nothing tying the list of links to an identity, and a model that has fetched several files has no way to attribute them.

The links are the entire point. An llms.txt with no links is a file that indexes nothing — it is a signpost with no destinations written on it.

Enough content to be meaningful. A file under 50 characters cannot carry a title, a summary and a link list. In practice a body this short is a placeholder someone committed and forgot, or a template that was never filled in.

A caveat worth being honest about

llms.txt is a community proposal from 2024, not a ratified standard. No major AI provider has publicly committed to reading it, and nobody can currently demonstrate that publishing one improves how your site is represented in AI answers.

That is why this finding is graded STANDARD rather than higher. It is worth fixing — you have already done the hard part by publishing the file, and making it valid is usually a five-minute edit — but it is not an emergency, and it should not be prioritised above indexing or crawlability work.

How to fix it

Edit llms.txt at the root of your site so it follows the llmstxt.org format. A minimal valid file looks like this:

## Your Site Name

> One sentence describing what your site is and who it serves.

### Main pages
- [Products](https://yoursite.com/products): What you sell
- [Pricing](https://yoursite.com/pricing): Plans and costs
- [Documentation](https://yoursite.com/docs): How to use the product

### Company
- [About](https://yoursite.com/about): Who we are
- [Contact](https://yoursite.com/contact): How to reach us

Fixing each violation

Missing H1 title — add a single # heading as the first line. One # only; ## is a section, not a title.

No links — add at least one entry in - [Label](https://full-url): description form. Use absolute URLs. Relative links (/about) do work and PixyScan resolves them, but absolute ones are unambiguous for any consumer.

Too short — the file needs real content. If you published a placeholder, replace it with your actual top pages, or remove the file entirely until you are ready to write one. An absent file is a cleaner state than a stub.

Worth doing while you are in there

Neither of these raises this issue, but both are recommended by the spec and PixyScan reports them alongside it:

  • Add a > summary blockquote under the title. One line of context before a model follows any link.
  • Group links under ## Section headings. This tells a model which links are documentation, which are products, and which are optional reading.

Choosing what to list

llms.txt is a curated index, not a sitemap. A short, well-chosen list of ten pages is more useful than an exhaustive list of four hundred. Include the pages you would want an assistant to cite when answering a question about your business; leave out pagination, tag archives and login pages.

Examples

1. No H1 title

The file opens with prose instead of a # Site name heading, so nothing identifies whose index this is.

Problem — raises missing_h1:

Acme Shoes

> We have sold running shoes online since 2011.

### Products

- [Running shoes](https://acme.com/running): Our main range

Fixed — the first line is an H1:

## Acme Shoes

> We have sold running shoes online since 2011.

### Products

- [Running shoes](https://acme.com/running): Our main range

The file describes the site but indexes nothing, so an assistant reading it is pointed at no pages.

Problem — raises no_links. The [Getting started](/docs/start) inside the fence does not count, because fenced code is excluded before parsing:

## Acme Shoes

> We have sold running shoes online since 2011.

Our documentation uses entries in this form:

```markdown
- [Getting started](/docs/start): Read this first
```

Fixed — real entries outside the fence:

## Acme Shoes

> We have sold running shoes online since 2011.

### Products

- [Running shoes](https://acme.com/running): Our main range
- [Trail shoes](https://acme.com/trail): For off-road

3. Plain URLs — reported as advice, not as a fault

A list of bare URLs is a working index. It is not the Markdown link form llmstxt.org specifies, so it is worth improving, but it does not raise this issue.

Passes, with the links_not_markdown recommendation:

## Acme Shoes

### Products

- Running shoes: https://acme.com/running
- Trail shoes: https://acme.com/trail

Better — each entry now carries a name an assistant can choose by:

## Acme Shoes

### Products

- [Running shoes](https://acme.com/running): Our main range
- [Trail shoes](https://acme.com/trail): For off-road

How PixyScan detects this

The file is fetched once per scan by the GEO & AI Engine Signals toggle group, which requests https://<origin>/llms.txt for issue #22. That fetch now keeps the response body, and this check reads it — previously the body was discarded as soon as the file was confirmed to exist, which is why a file containing a single junk word scored a clean pass.

Parsing and validation are performed by validateLlmsTxtFormat in apps/crawler/utils/llmsTxt.js; the finding is raised by checkLlmsTxtQuality in apps/crawler/seo-audit-checks.js.

Detection steps

  1. Precondition. The check runs only when issue #22 found a real file. If llms.txt is missing, or answered 200 with an HTML soft-404, this check is skipped entirely — reporting the same absence under two issue numbers would be double-counting.

  2. Parse. The body is read as Markdown, up to a 2 MB ceiling:

    • the first # line becomes the title
    • the first > line becomes the summary
    • each ## line opens a section
    • every [Label](url) match is collected as a link, along with any : description that follows it
    • a list item that is a plain URL — - https://site.com/docs or - Home: https://site.com/ — is also collected as a link, with the text before the colon as its label. Restricted to list items on purpose: a URL mentioned in prose, such as a Website: https://… line in a contact block, is not an index entry.
    • when the file contains no ## at all, # headings after the title are read as sections

    The last two rules exist because of what real files do. In a survey of eleven live llms.txt files, two wrote every entry as a plain URL — one of them a 206-entry documentation index — and reading only Markdown links scored both as having no links whatsoever.

  3. Fenced code is excluded. Content inside ``` or ~~~ fences is blanked before parsing. Files that document an API often contain example Markdown links inside a fence; counting those would let a file whose real index is empty report links, which is exactly the defect this check exists to catch.

  4. Judge, in two tiers. Violations raise this issue: missing_h1, no_links (no page URLs of any syntax), too_short (body under 50 characters after trimming). Recommendations are returned in the finding's details but never raise it on their own: missing_summary, no_sections, links_not_markdown, sections_use_h1, oversized.

    no_links is deliberately about substance rather than syntax. It used to mean "no Markdown links", which made it fire on files carrying eight and 206 working URLs respectively — telling sites with functioning indexes that they pointed an AI at nothing. Syntax is now advice (links_not_markdown), and the violation is reserved for a file that genuinely lists no pages.

What gets stored

The finding's details carry the parsed title, linkCount and sectionCount, the list of violations, and the list of recommendations. Each entry has a stable code and a human-readable message.

Known limits

  • The file is read as served. PixyScan does not execute JavaScript, so a llms.txt generated client-side is not visible to this check — though such a file would not be visible to an AI crawler either.
  • Only the first 2 MB is parsed. An llms.txt larger than that is reported with the oversized recommendation. The format is meant for an index; bulk content belongs in llms-full.txt (issue #19).

What we store

Storage Level

Site Level — the file is fetched once per scan and graded once, so the finding carries no URL.


Database Table / Prisma Model

AuditIssue

The format verdict has no dedicated audit table. llms.txt is a single site-wide file, and everything this check concludes about it is carried on the finding itself. The neighbouring flags for whether the file exists (#19, #22) live on SiteCrawlBehaviourData.aiEngineConfig; this issue never writes there.


Fields Used

Field Type Description
scanId String The scan the finding belongs to
urlId String? Always null — this is a site-level finding
issueCode String llms_txt_malformed
details Json? message, violations[] (each {code, message}), recommendations[] (each {code, message}), title (the parsed H1), linkCount, sectionCount
createdAt DateTime When the finding was written

violations is what raised the issue: missing_h1, no_links, too_short. recommendations is advice carried for display and never raises it: missing_summary, no_sections, links_not_markdown, sections_use_h1, oversized. The two are kept as separate arrays so the report can show both without the reader having to guess which ones are faults.


Detection Dependencies

  • llms.txt — the response body, retained from the fetch that answers issue #22
  • HTTP Response — the file must have answered 200 with a real body; a soft-404 skips this check entirely

Further reading