Skip to content
Issue docs

llms.txt links to pages that do not resolve

Standardllms_txt_broken_linksIssue 24

What is this issue?

One or more of the links listed in your llms.txt does not resolve. An AI assistant that follows your own curated index lands on an error page.

llms.txt is a hand-maintained file. Unlike an XML sitemap, nothing generates it from your live routes, so it drifts: a page gets renamed, a section is retired, a URL gains a locale prefix — and the index still points at the old address. Because almost nothing else reads the file, the breakage is invisible until something follows it.

This finding lists the specific links that returned an error, with the HTTP status each one produced.

What counts as broken

Only a link PixyScan confirmed to be dead. A link is reported here when it was requested and the server answered with a 4xx or 5xx status.

Three things are deliberately not reported as broken:

  • Links to pages PixyScan did not crawl. Absence from the crawl proves nothing — the crawl is a sample. Such links are fetched directly before any verdict is reached.
  • Links that could not be reached. A timeout, DNS failure or TLS error means the question could not be asked. That is recorded as unknown, never as broken.
  • Links to other domains. A dead link on a third party's site is not your defect, and PixyScan will not generate traffic against a site it was not asked to scan.

Why it matters

The whole value of llms.txt is that it is curated. You are telling an AI assistant: skip the guesswork, these are the pages that matter. A broken entry inverts that — it routes the assistant, with your explicit endorsement, to a page that does not exist.

This is worse than a broken link elsewhere on your site for two reasons.

The file exists to be followed. A dead link buried in a footer is a minor annoyance. A dead link in a file whose only purpose is to be a list of links is a failure of the file's single job.

Nobody notices. Your ordinary broken links get caught by crawlers, by analytics, by users complaining. Almost nothing reads llms.txt, so an entry can rot there for a year. This check is likely the only thing looking.

There is also a credibility dimension. If an assistant follows two entries and both 404, it has no reason to trust the third — and the file you published to improve how your site is represented has made it worse.

Graded STANDARD

Like every llms.txt finding, this is graded STANDARD. llms.txt is a community proposal that no AI provider has publicly committed to consuming, so a defect in it cannot outrank a defect in something search engines demonstrably read. It is, however, one of the cheapest things on this list to fix.

How to fix it

Open the finding's details, which list each broken URL alongside the HTTP status it returned. For each one:

404 — the page is gone. Either update the entry to the page's new address, or remove the entry. If the content moved, point at the new URL directly rather than at a URL that redirects — an index should name its destination.

403 or 401 — the page requires authentication. It does not belong in llms.txt. The file lists pages you want an AI assistant to read; a page it cannot open is noise.

5xx — the page is erroring. This is a bug on the page itself, not in llms.txt. Fix the page; the entry is fine.

Preventing the drift

llms.txt is hand-maintained, so it will drift again unless something keeps it honest. Two options, in order of effort:

  1. Generate it. If your site has a content source — a CMS, a docs framework, a route manifest — emit llms.txt from it at build time. A generated file cannot point at a page that no longer exists.

  2. Check it in CI. A short script that reads llms.txt and requests each URL will catch a broken entry at the commit that caused it, rather than at the next scan.

If neither is practical, keep the file short. A ten-entry index that is correct is worth more than a hundred-entry index that is half stale — and it is far easier to review by eye when you rename something.

Examples

1. A renamed page left in the index

The /pricing page was moved to /plans and the file was never updated.

Problem — the probe answers 404 and the link is reported:

### Company

- [Pricing](https://acme.com/pricing): Plans and costs

Reported as:

url: https://acme.com/pricing
label: Pricing
httpStatus: 404

Fixed — point at the live address:

### Company

- [Plans](https://acme.com/plans): Plans and costs

2. A locale prefix added site-wide

The site moved every page under /en/, and the index still lists the old flat paths.

Problem — every entry answers 404:

### Documentation

- [Getting started](https://acme.com/docs/start)
- [API reference](https://acme.com/docs/api)

Fixed:

### Documentation

- [Getting started](https://acme.com/en/docs/start)
- [API reference](https://acme.com/en/docs/api)

None of these raise the issue, and it is worth knowing why — each one is a case where the check cannot honestly say the link is dead.

Link What happened Verdict
https://acme.com/guide HEAD answered 404, GET answered 200 ok — the retry is what makes this correct
https://acme.com/admin/settings disallowed by robots.txt, never requested unchecked, with a reason
https://partner.example/tools another domain not verified at all
https://acme.com/slow-report the request timed out unknown, counted in uncheckedCount

The 404-then-GET retry exists because some servers answer HEAD with 404 for pages they serve perfectly well. Reporting one of those would be a confident accusation about a page that is fine.

How PixyScan detects this

Verification is performed by verifyLlmsTxtLinks in apps/crawler/seo-audit-checks.js, once per scan, after the crawl has finished.

The design constraint behind almost every rule below is that the crawl is a sample. On a site with 10,000 pages and a 500-page budget, most of the links in an llms.txt will be absent from the crawled set simply because they were never requested. Concluding "this link is broken" from that absence would be wrong in the most damaging way possible — a confident accusation about a page that is perfectly fine. So absence never decides anything; it only triggers a request.

Detection steps

  1. Partition the links. Entries are resolved against the site's origin and split into same-site, off-domain and unusable (mailto:, tel:, and anything that is not a web URL). Only same-site links are verified. Duplicate spellings of one page collapse to a single entry.

  2. Skip what the crawl already proved. A link whose canonical key matches a page the crawl successfully fetched is recorded as good with no request at all. On a well-covered site this is most of the file.

  3. Skip what we were told not to fetch. A link disallowed by robots.txt, or matching the site's exclude patterns, is not requested — and cannot be called broken without being requested. It is recorded as unchecked, with a reason.

  4. Probe the rest. HEAD first, then GET when the server answers 405, 501 or 404. The 404 retry is deliberate: some servers answer HEAD with 404 for pages they serve perfectly well on GET, and "your llms.txt links to a dead page" about a live page is the most damaging thing this check could say.

  5. Classify. 2xx/3xx is ok. 4xx/5xx is broken. A timeout, DNS failure or TLS error is unknown — we could not ask the question, and our own network trouble must not appear in a customer's report as their defect.

  6. Raise. The issue fires only when at least one link is broken.

Caps

At most 25 links are probed per scan, at a concurrency of 4, spaced by the crawl's own politeness delay. This runs after the crawl has finished and its rate limiting has wound down, so an uncapped check on a 500-link file would mean 500 requests fired outside the polite window the crawl spent its whole run honouring.

Links past the cap are reported in linksNotChecked with reason over-check-limit. They are never silently dropped — silent truncation would read as "we checked everything".

What gets stored

brokenCount and a brokenLinks array (URL, label, HTTP status), plus checkedCount and uncheckedCount so that "3 broken" is never mistaken for "and the rest are fine".

What we store

Storage Level

Site Level — one file is verified once per scan, so the finding carries no URL. The dead links themselves are listed inside it.


Database Table / Prisma Model

AuditIssue

Link verification has no dedicated audit table. The links come from a single site-wide file rather than from any crawled page, so there is no per-URL row to hang them on, and the whole result is carried on the finding.


Fields Used

Field Type Description
scanId String The scan the finding belongs to
urlId String? Always null — this is a site-level finding
issueCode String llms_txt_broken_links
details Json? message, brokenCount, brokenLinks[], checkedCount, uncheckedCount
createdAt DateTime When the finding was written

brokenLinks holds one entry per confirmed dead link: url (as resolved against the site origin), label (the link text from the file) and httpStatus (the code the server answered, or null).

checkedCount and uncheckedCount are stored alongside the count on purpose. Without them "3 broken" reads as "and the rest are fine", which is not what a capped, policy-filtered check can claim. uncheckedCount covers links skipped by robots.txt or an exclude pattern, links past the 25-link probe cap, and links whose request failed outright.


Detection Dependencies

  • llms.txt — the link list being verified
  • HTTP Response — a HEAD probe per unverified link, retried as GET on 404, 405 or 501
  • robots.txt — a disallowed link is never requested, so it is never called broken
  • Crawl results — links the crawl already fetched successfully are passed without a second request
  • Site exclude patterns — the other half of "URLs we were told not to fetch"

Further reading