Skip to content
Issue docs

Canonical target canonicalises to another page

Standardcanonical_chainIssue 162

What is this issue?

This issue reports a page whose <link rel="canonical"> names a page that, in turn, names a different canonical of its own. Page A says "index B"; page B says "index C".

A canonical should name the end of the line — a page that is its own canonical.

For a page to pass this check:

  • The page its canonical names is its own canonical.

Example: /shoes?colour=red → canonical /shoes?page=1, and /shoes?page=1 → canonical /shoes.

A loop — A names B and B names A — is reported the same way, and the finding says it is a loop.

Why it matters

  • Each hop is a chance to stop. Google usually follows one step of a canonical chain, but canonicals are hints, and every extra hop makes it likelier your chosen page is not the one indexed.
  • A loop gives no answer at all. Two pages naming each other leave search engines to pick one themselves.
  • It usually means two templates disagree about which address is canonical — for example a filter template and a pagination template.

How to fix it

  1. Point the canonical straight at the end of the chain. If A → B → C, make A name C.
  2. For a loop, decide which of the two pages is the canonical version and make both name it.
  3. Make templates agree. If filtered, sorted and paginated views each set their own canonical, settle one rule — usually "the unfiltered first page" or "itself" — and apply it everywhere.

Examples

Example 1: One extra hop (reported)

<!-- https://example.com/shoes?colour=red -->
<link rel="canonical" href="https://example.com/shoes?page=1">

<!-- https://example.com/shoes?page=1 -->
<link rel="canonical" href="https://example.com/shoes">

Corrected:

<!-- https://example.com/shoes?colour=red -->
<link rel="canonical" href="https://example.com/shoes">

Example 2: A loop (reported, with loop: true)

/a → canonical /b, and /b → canonical /a.

Example 3: The target names itself in another spelling (not reported)

/b → canonical https://example.com/b/#main. That is still /b.

Example 4: The "chain" is a redirect (reported as a redirect instead)

/a → canonical /old, which answers 301. That is Canonical points at a redirect, and is never reported here as well.

How PixyScan detects this

  1. Reads the canonical of every page that answered 2xx, resolves it and puts it in the stored address form.
  2. Looks the target up among the pages this scan fetched.
  3. Reads the target's own canonical and resolves it the same way.
  4. Reports the page when the target answered 2xx and its canonical names a different address. The finding names that address, and flags a loop when it is the reported page itself. A target canonical that is not an address at all is not followed.

What is never judged

  • A self-referencing canonical. A page naming itself — in any spelling: a trailing slash, a fragment, a tracking parameter — is the correct and common case and is not looked up.
  • A canonical whose target this crawl did not fetch. Another domain, an address no link led to, a sitemap entry the scan's URL allowance would not stretch to, or one your exclusion rules dropped. PixyScan has no measured answer for it and does not guess. Every finding records how many canonicals were judged and how many pointed at addresses the crawl did not fetch, so a result can always be read against how much of the site it covers.
  • A canonical that is not an address. Prose in the href, or a non-http scheme. That is the malformed-canonical finding (#63).
  • A page that did not itself answer 2xx. An error page or a redirect serves no document a search engine would read a canonical from.

How the four canonical-target checks fit together

Each page's canonical is judged once, in this order:

  1. The target answered 4xx or 5xx → Canonical points at an error page (#159), and nothing else.

  2. The target redirects → Canonical points at a redirect (#160), and nothing else. In particular it is never also reported as a canonical chain: when the "chain" is a redirect, the redirect is the finding.

  3. The target answered 2xx:

    • it carries noindex → Canonical points at a noindex page (#161);
    • its own canonical names another address → Canonical target canonicalises to another page (#162).

    These two can both be reported for one page, because they are different faults on the target with different fixes.

What we store

Storage Level

Page Level — reported against the page that declares the canonical.


Database Table / Prisma Model

Read from PageSeoBasicsData (both pages), Url (the target's crawl state and status) and RedirectChain (where a redirect lands). The finding is stored on audit_issues.details.


Fields Used

Field Type Description
page_seo_basics_data.canonical_url String? The canonical declared on the reported page, and on its target
page_seo_basics_data.indexability_noindex_absent Boolean? false when the target carries noindex in a robots meta tag or an X-Robots-Tag header; null is never read as noindex
urls.url String The stored address the canonical is matched to, after resolving and normalising it
urls.status Enum Only completed or failed rows count as fetched; pending, skipped and excluded are not judged
urls.status_code Int? What the target answered. Status 0 (our own failed request) is not an answer and is not judged
redirect_chains.original_url / final_url String Where a target that redirects lands, for redirects recorded under the address that was requested

Stored Fields on the finding

Field Type Description
message / recommendation String What was found at the target, and what to change
url String The page being reported
canonicalUrl String The canonical as the page declared it
targetUrl / targetUrlId String The page the canonical resolved to
targetStatusCode Int What the target answered
targetCanonicalUrl String The address the target canonicalises to
loop Boolean True when the target canonicalises straight back at the reported page
scope String Always targets-fetched-in-this-scan
basis String A sentence stating how many canonicals were judged and how many were not
pagesWithCanonical, selfCanonical, unresolvable, targetsNotFetched, targetsJudged Int The coverage counts behind basis

Detection Dependencies

  • HTML Document — the canonical tag on both pages, and the target's robots meta tag
  • HTTP Response — the target's status code and X-Robots-Tag header
  • Redirects — the stored redirect chain, when the target redirects

Further reading