Skip to content
Issue docs

Canonical points at an error page

Importantcanonical_to_error_pageIssue 159

What is this issue?

This issue reports a page whose <link rel="canonical"> names an address that answers with an error — a 4xx such as 404 or 410, or a 5xx server error.

A canonical says "index that page instead of this one". When that page is gone or broken, the instruction points at nothing.

For a page to pass this check:

  • The page its canonical names answers HTTP 200 (or another 2xx).

Example: /shoes?colour=red declares <link rel="canonical" href="https://example.com/shoes-old">, and /shoes-old was deleted and now answers 404.

A page whose canonical names itself is never reported here, and neither is one whose canonical names an address this scan did not fetch.

Why it matters

  • The canonical is ignored. Google does not consolidate a page into an address that returns an error, so the duplicate this canonical was meant to fold away is indexed — or dropped — on its own terms.
  • Ranking signals are lost. Links pointing at the page were meant to count towards the canonical. With the canonical broken they count towards neither.
  • It usually means a template is stale. Canonicals are generated, so one broken target is often many pages pointing at the same removed address.
  • AI search and answer engines read canonicals to decide which version of a page to quote. A canonical to an error gives them nothing to quote.

How to fix it

  1. Point the canonical at the live page. Usually that is the page itself — a self-referencing canonical — or the current address of the content.
  2. If the target moved, use its new address directly. Do not rely on a redirect from the old one (that is a different finding).
  3. If the target was removed on purpose, the pages pointing at it need their own canonical: themselves, or the closest page that still exists.
  4. Fix it in the template. If many pages share the broken target, the canonical is generated from one place — correct it there.
  5. Re-scan and confirm the target answers 200.

Examples

Example 1: Canonical to a deleted page (reported)

<!-- https://example.com/shoes?colour=red -->
<link rel="canonical" href="https://example.com/shoes-old">

https://example.com/shoes-old → 404 Not Found

Corrected:

<link rel="canonical" href="https://example.com/shoes">

Example 2: Canonical to a page with a server error (reported)

/blog/post?print=1 → canonical /blog/post, which answered 503 during the scan. Fix the server error; if the page is healthy now, the next scan clears the finding.

Example 3: Self-referencing canonical (not reported)

<!-- https://example.com/shoes -->
<link rel="canonical" href="https://example.com/shoes/">

The trailing slash does not make it a different page.

Example 4: Canonical to another domain (not judged)

<link rel="canonical" href="https://partner.example.net/shoes"> — the scan did not fetch the partner site, so PixyScan has nothing to judge it on.

How PixyScan detects this

This is decided after the crawl, by joining each page's canonical to what the crawl found at that address.

  1. Reads the canonical of every page that answered 2xx.
  2. Resolves it and puts it in the stored form — relative paths resolved against the page, fragment, trailing slash and tracking parameters dropped, query parameters sorted — so it names the same row the crawl filed the target under.
  3. Looks the target up among the pages this scan fetched.
  4. Reports the page when the target answered 400 or above.

What is never judged

  • A self-referencing canonical. A page naming itself — in any spelling: a trailing slash, a fragment, a tracking parameter — is the correct and common case and is not looked up.
  • A canonical whose target this crawl did not fetch. Another domain, an address no link led to, a sitemap entry the scan's URL allowance would not stretch to, or one your exclusion rules dropped. PixyScan has no measured answer for it and does not guess. Every finding records how many canonicals were judged and how many pointed at addresses the crawl did not fetch, so a result can always be read against how much of the site it covers.
  • A canonical that is not an address. Prose in the href, or a non-http scheme. That is the malformed-canonical finding (#63).
  • A page that did not itself answer 2xx. An error page or a redirect serves no document a search engine would read a canonical from.

How the four canonical-target checks fit together

Each page's canonical is judged once, in this order:

  1. The target answered 4xx or 5xx → Canonical points at an error page (#159), and nothing else.

  2. The target redirects → Canonical points at a redirect (#160), and nothing else. In particular it is never also reported as a canonical chain: when the "chain" is a redirect, the redirect is the finding.

  3. The target answered 2xx:

    • it carries noindex → Canonical points at a noindex page (#161);
    • its own canonical names another address → Canonical target canonicalises to another page (#162).

    These two can both be reported for one page, because they are different faults on the target with different fixes.

What we store

Storage Level

Page Level — reported against the page that declares the canonical.


Database Table / Prisma Model

Read from PageSeoBasicsData (both pages), Url (the target's crawl state and status) and RedirectChain (where a redirect lands). The finding is stored on audit_issues.details.


Fields Used

Field Type Description
page_seo_basics_data.canonical_url String? The canonical declared on the reported page, and on its target
page_seo_basics_data.indexability_noindex_absent Boolean? false when the target carries noindex in a robots meta tag or an X-Robots-Tag header; null is never read as noindex
urls.url String The stored address the canonical is matched to, after resolving and normalising it
urls.status Enum Only completed or failed rows count as fetched; pending, skipped and excluded are not judged
urls.status_code Int? What the target answered. Status 0 (our own failed request) is not an answer and is not judged
redirect_chains.original_url / final_url String Where a target that redirects lands, for redirects recorded under the address that was requested

Stored Fields on the finding

Field Type Description
message / recommendation String What was found at the target, and what to change
url String The page being reported
canonicalUrl String The canonical as the page declared it
targetUrl / targetUrlId String The page the canonical resolved to
targetStatusCode Int What the target answered
scope String Always targets-fetched-in-this-scan
basis String A sentence stating how many canonicals were judged and how many were not
pagesWithCanonical, selfCanonical, unresolvable, targetsNotFetched, targetsJudged Int The coverage counts behind basis

Detection Dependencies

  • HTML Document — the canonical tag on both pages, and the target's robots meta tag
  • HTTP Response — the target's status code and X-Robots-Tag header
  • Redirects — the stored redirect chain, when the target redirects

Further reading