Crawled pages are missing from llms.txt
What is this issue?
This is a suggestion, not a defect. It does not count against your health score.
PixyScan crawled your site and found pages that your llms.txt does not mention. This report lists them so you can decide whether any of them belong in your index.
Why this is advice and not a fault
llms.txt is a curated index, not a sitemap. Leaving a page out is frequently the correct decision — pagination, tag archives, thank-you pages, legal boilerplate and duplicate landing pages all belong outside it. A deliberately short llms.txt is often a better one.
So PixyScan can tell you what is absent. It cannot tell you that the absence is wrong. That judgement is editorial and it is yours, which is why this carries no severity.
How the list is ordered
The pages are ordered by crawl depth — how many links from the homepage PixyScan had to follow to reach each one — with the closest first.
This is a proxy, and it is labelled as one. PixyScan has no traffic data, so it cannot rank your pages by importance. Depth is a reasonable stand-in because sites tend to link their significant pages closer to the front door, but it is not a claim about value. A confident wrong ordering would be worse than an explicitly arbitrary one.
Scope of the numbers
Every figure here describes the pages this scan actually crawled, not your whole site. If the scan hit its page limit, the message says so explicitly. A scan that saw 500 of your 10,000 pages can report on 500.
Why it matters
Most llms.txt files are written once, when the idea is new, and then never revisited. The site grows around them. A file that listed the right eight pages in 2024 lists eight of your forty pages now, and the thirty-two it omits include the ones you have since built your business on.
This report is the periodic nudge to look. It is the one llms.txt check that requires actually crawling the site, which is why very little else can produce it: a validator can tell you the file parses, but only a crawler can tell you whether it still describes the site it sits on.
What to do with it
Read the list and ask, for each page: if an AI assistant were answering a question about my business, would I want it to have found this?
Usually the answer is no for most of the list and yes for two or three. Those two or three are the value of this report.
What not to do with it
Do not add everything. An llms.txt listing every URL is a sitemap, and you already have one of those. The file's usefulness comes from being selective; a hundred-entry index dilutes the signal that made the ten-entry version worth reading.
Do not treat a high count as a failure. A site with 400 pages and a 12-entry llms.txt will report 388 pages absent. That is not a problem — it may be exactly right.
How to fix it
There is nothing here that must be fixed. This is a prompt to review, not a defect to clear.
Reviewing the list
Work through the reported pages and sort them into three piles:
Belongs in the index. Pages an AI assistant should cite when answering a question about you: products, pricing, core documentation, key service pages, substantial guides. Add these.
Deliberately excluded. Pagination, tag and category archives, search results, login and account pages, thank-you and confirmation pages, near-duplicate landing pages, legal boilerplate. Leave these out.
Should not be crawlable either. Occasionally this list surfaces a page that should not be indexed at all. That is a robots.txt or noindex question rather than an llms.txt one, but it is worth acting on when it appears.
Adding an entry
Put it under a ## section that already exists, or open a new one:
### Guides
- [How to migrate](https://yoursite.com/guides/migrate): Step-by-step migration walkthroughWrite the description for a reader who has not seen the page. "Step-by-step migration walkthrough" earns its place; "Migrate" does not.
Making it stay current
If this report keeps returning a long list, the file is being maintained by hand and losing. Generating llms.txt from your CMS or route manifest at build time — filtered to the page types you actually want indexed — turns this from a recurring chore into a one-off.
Examples
1. Pages worth adding
The crawl reached 42 pages. The index lists six, and the three shallowest omissions are real product pages.
Reported:
missingCount: 36
sample:
- https://acme.com/trail (depth 1)
- https://acme.com/warranty (depth 1)
- https://acme.com/size-guide (depth 2)Acted on — the entries that belong in a curated index are added:
## Acme Shoes
### Products
- [Running shoes](https://acme.com/running): Our main range
- [Trail shoes](https://acme.com/trail): For off-road
### Help
- [Warranty](https://acme.com/warranty): What is covered, and for how long
- [Size guide](https://acme.com/size-guide): How our sizes run2. Pages correctly left out — no action needed
The same report also lists pages that have no business being in an index for an assistant:
- https://acme.com/blog/page/7 (depth 3)
- https://acme.com/tag/waterproof (depth 3)
- https://acme.com/checkout/thanks (depth 4)Pagination, tag archives and a post-purchase confirmation are all reasonable omissions. A shorter, deliberately curated llms.txt is often the better file, which is why this carries no severity — the report tells you what is absent, and whether that is wrong is your call.
3. A spelling difference that is not reported
The file writes a relative path; the crawler recorded the canonical absolute address.
In llms.txt:
- [About us](/about)Recorded by the crawl:
https://www.acme.com/about/These match. Both sides are folded to the same key first — scheme, a leading www., the default port, a trailing slash, the fragment and query-parameter order all collapse — so the page is counted in matchedCount, not reported as missing. Without that folding a site would be told every link it publishes is absent from its own crawl. Path case is deliberately not folded, because paths are case-sensitive on most servers.
How PixyScan detects this
The comparison is performed by compareCoverage in apps/crawler/utils/llmsTxt.js, called from checkLlmsTxtQuality once per scan after the crawl has finished. It needs the crawl's results, which is why the GEO & AI Engine Signals group receives the crawled URL set from the site-level context.
Detection steps
Build the listed set. Same-site links from
llms.txtare reduced to canonical keys. Off-domain and non-web links are excluded.Build the crawled set. Every URL the crawl reached, minus anything
robots.txtdisallows and anything matching the site's exclude patterns.Subtract. Crawled pages whose key is not in the listed set are the reported pages.
Order by depth, shallowest first, with unrecorded depth last.
Report, but only when the crawl saw at least 2 eligible pages. The sample is capped at 20 entries, with
truncatedSampleset when more were found.The floor used to be five, and it quietly excluded exactly the sites this suggestion is easiest to act on: a brochure site, a docs root, a landing page with three sections. A four-page site that lists one page in its
llms.txthas three pages missing from its index and every reason to hear about it — and because the finding names the pages, the number never stands on its own. A one-page crawl is still excluded: there the count adds nothing to what issues #22 and #23 have already said about the file.
URL matching
Both sides are reduced with canonicalUrlKey, which folds the spellings that address the same page: scheme (http/https), a leading www., default ports, a trailing slash, the fragment, and query-parameter order. Relative links in llms.txt are resolved against the site origin.
Without this the report would be nonsense in a very visible way — a site whose file says /about and whose crawler recorded https://www.example.com/about/ would be told that every link it publishes is missing. Path case is not folded, because paths are case-sensitive on most servers.
Three deliberate restraints
The comparison runs in one direction only. A page listed in llms.txt that PixyScan did not crawl is never reported here. The crawl is a sample, and absence from it usually means the page was never requested rather than that it is gone. Judging those links belongs to issue #24, which fetches them.
Excluded and disallowed pages are removed first. They are absent by instruction. Counting them as "missing from llms.txt" would report PixyScan's own configuration back as the customer's defect.
A capped crawl changes the wording. When the scan stops at its page budget, the finding says "N of the M pages crawled in this scan", and sets partialCrawl: true. "17 pages are not listed" is a claim about the whole site, and it is only supportable when the whole site was seen.
What gets stored
crawledCount, listedCount, missingCount, matchedCount, the capped sample (URL and depth), truncatedSample, partialCrawl, and sampleOrdering: "closest-to-homepage-first" — named so the ordering is never mistaken for a ranking by importance.
What we store
Storage Level
Site Level — the comparison is between one file and the whole crawled set, so the result belongs to the scan rather than to any single URL.
Database Table / Prisma Model
AuditIssue
This is a suggestion rather than a defect: it carries no severity and does not reduce the health score. It has no dedicated audit table — the pages it names are already Url rows from the crawl, and what is stored here is the comparison, not the pages.
Fields Used
| Field | Type | Description |
|---|---|---|
| scanId | String | The scan the suggestion belongs to |
| urlId | String? | Always null — this is a site-level finding |
| issueCode | String | llms_txt_coverage |
| details | Json? | The comparison result, listed below |
| createdAt | DateTime | When the finding was written |
What details carries:
| Key | Type | Description |
|---|---|---|
| message | String | Worded differently when the crawl was capped, so the claim matches its evidence |
| crawledCount | Int | Eligible pages the crawl reached, after removing disallowed and excluded URLs |
| listedCount | Int | Distinct same-site pages the file lists |
| missingCount | Int | Crawled pages the file does not mention |
| matchedCount | Int | Crawled pages the file does mention |
| sample | Json | Up to 20 missing pages, each {url, depth}, shallowest first |
| truncatedSample | Boolean | Whether more were found than the sample shows |
| partialCrawl | Boolean | Whether the crawl stopped at its page budget |
| sampleOrdering | String | Always closest-to-homepage-first |
sampleOrdering is stored as a literal string rather than left implicit. The order is crawl depth, which is a fact about link distance from the homepage — naming it prevents the list from being read as a ranking by importance, which there is no traffic data to make.
Detection Dependencies
- llms.txt — the listed set
- Crawl results — the crawled set, and each page's crawl depth for the ordering
- robots.txt — disallowed pages are removed before the comparison
- Site exclude patterns — excluded pages are removed before the comparison
- Crawl budget state — whether the scan hit its page limit, which changes what the finding claims