Skip to content
Issue docs

Broken internal links and pages

Criticalno_broken_internalIssue 8

What is this issue?

A page on your own site answered with a client error — an HTTP status in the 400 range — when PixyScan requested it, usually because it was removed, renamed or never existed. Every link to it on your site is a broken link.

The common ones:

Status What the server is saying
404 Not Found There is nothing at this address
410 Gone There was something here and it was removed on purpose
401 / 403 The page exists but refuses this visitor
429 Too Many Requests The server is refusing requests from this client for now

For a page to pass this check:

  • Every internal page PixyScan fetched, or found linked from another page, answers without a 4xx status.

Example: your navigation links to /about-us, but the page was renamed to /about and /about-us now answers 404 Not Found.

Server errors are a separate check. A page answering 5xx (500, 502, 503, 504) is reported as "Internal pages returning a server error (5xx)" (#165): the page should exist and the server failed to produce it, which is a server fix rather than a content or link fix. One page is only ever reported under one of the two.

Why it matters

  • Readers hit a dead end. Every visitor who follows the link gets an error page instead of what they clicked for, and many leave.

  • Search engines drop the page. A URL that answers 404 or 410 is removed from the index; whatever it ranked for is lost.

  • Crawl budget is wasted. Crawlers keep requesting addresses your own links point at, even when they lead nowhere.

  • Link signals are thrown away. Internal links pass standing between your pages. A link to a missing page passes it to nothing.

  • It signals neglect. A site with many broken links reads as unmaintained, to readers and to quality reviewers.

Effect on the health score

This is a critical issue under the Link Integrity lens.

How to fix it

  1. Find who links to it. The finding lists the pages linking to the broken address. If many pages do, the link is in a shared header, footer or navigation component — fix it once there.

  2. Decide what the address should do:

    • The page moved: add a 301 redirect from the old address to the new one, and update the links to point at the new address directly.
    • The page was removed on purpose: remove or replace the links, and let the address answer 404 or 410.
    • The address is a typo: correct the link.
  3. For 401 / 403: if the page should be public, fix the permission or the firewall rule; if it is private, stop linking to it from public pages.

  4. For 429: the server is rate limiting the crawler. Allow well-behaved crawlers through, or lower the scan's request rate in its settings.

  5. Update your CMS links when pages move. Many CMSs can create redirects automatically when a page's address changes.

  6. Re-scan to confirm every internal page answers 200.

Examples

Example 1: A renamed page

Fails: the homepage links to a page that was renamed.

<a href="/about-us">About us</a>
GET /about-us → 404 Not Found

Passes: the link points at the new address, and the old one redirects.

<a href="/about">About us</a>
GET /about-us → 301 Location: /about
GET /about    → 200 OK

Example 2: A page removed on purpose

Fails: an old campaign page answers 410 Gone, and the footer still links to it.

Passes: the footer link is removed. The address can keep answering 410.

Example 3: A server error — reported elsewhere

GET /checkout → 500 Internal Server Error

Not reported here. A 5xx is reported under "Internal pages returning a server error (5xx)" (#165).

How PixyScan detects this

  1. When the crawler fetches a page itself. Every page PixyScan requests during the crawl has its HTTP status recorded. A page that answers 400–499 is recorded as failed with its real status, is not audited further (an error page has no title or headings worth checking), and is raised under this check. A page answering 500–599 is raised under #165 instead.

  2. After the crawl, for pages other pages link to. PixyScan reads the status of every internal link target. On plans that include link checks (Hobby and above), each distinct target is also requested once with a lightweight HEAD request (falling back to a one-byte GET if the server refuses HEAD). A target answering 4xx is raised here — once per page, however many pages link to it.

  3. One code per page. The boundary — 400 to 499 here, 500 to 599 under #165 — is the same rule at fetch time and after the crawl. If two passes saw the same page differently, PixyScan keeps the finding that matches the page's most recent status, so the page never carries both.

  4. What is not reported. A request that never got an answer — a timeout, a DNS failure, a refused connection — has no status and is not reported as broken; it is reported as "Page failed to load or timed out during the crawl" (#193). An address the crawl never requested (excluded, beyond the page limit) is not judged at all.

  5. Respects your lens settings. Nothing is raised on a scan with the Link Integrity lens switched off.

Note: PixyScan reads the HTML your server sends. Links that only exist after JavaScript runs are not followed.

What we store

Storage Level

Page Level


Database Table / Prisma Model

Url (urls) holds the status the finding is decided from. InternalLink (internal_links) is how a linked page is found after the crawl, and how the pages linking to it are listed. The finding is stored on audit_issues.details.


Fields Used

Field Type Description
urls.status_code Int The HTTP status the page answered with. 400–499 raises this check; 500–599 raises #165. Null means never measured and is not judged
urls.status Enum failed for a page the crawler fetched and found answering an error
internal_links.source_num Int The page carrying the link — listed as evidence on the finding
internal_links.target_num Int The linked page (by its urls.num)
internal_links.anchor_text String The words of each link to the broken page, listed as evidence

Stored Fields on the finding

Field Type Description
message String "Page returned HTTP 404." or the post-crawl wording for a linked page
url String The page that answered with the error
statusCode Int The 4xx status it answered with

Detection Dependencies

  • HTTP Response — the status of the page itself
  • Internal Links — which pages are linked, and from where
  • Plan — the post-crawl HEAD probe of link targets runs on Hobby and above; the fetch-time check runs on every plan

Further reading