Skip to content
Issue docs

Canonical points at a redirect

Importantcanonical_to_redirectIssue 160

What is this issue?

This issue reports a page whose <link rel="canonical"> names an address that redirects somewhere else.

The canonical should be the final address — the page that actually answers 200. Naming an address that redirects asks search engines to index one URL while the server sends them to another.

For a page to pass this check:

  • The page its canonical names answers directly, without redirecting.

Example: /shoes?colour=red declares <link rel="canonical" href="http://example.com/shoes">, and http:// redirects to https://example.com/shoes.

A trailing-slash-only redirect (/shoes → /shoes/) is the same page in PixyScan's storage and is not reported.

Why it matters

  • Two signals point in different directions. The canonical says "this address"; the redirect says "no, that one". Google treats both as hints and may honour neither, choosing a canonical of its own.
  • Consolidation is slower. Every recrawl has to follow the redirect to find the real page.
  • It often reveals a migration left half done — an http to https move, a www change, or a renamed section whose templates still emit the old address.

How to fix it

  1. Replace the canonical with the address the redirect ends on. That is the page that actually answers 200.
  2. Check the scheme and host. The most common cause is a canonical still written with http:// or the old host after a move.
  3. Fix it at the source. Canonicals are generated — correct the base URL setting or template, not page by page.
  4. Keep the redirect. It still serves old links; only the canonical needs to change.

Examples

Example 1: Canonical still on http (reported)

<!-- https://example.com/shoes?colour=red -->
<link rel="canonical" href="http://example.com/shoes">

http://example.com/shoes → 301 → https://example.com/shoes

Corrected:

<link rel="canonical" href="https://example.com/shoes">

Example 2: Canonical to a renamed page (reported)

/blog/old-title → 301 → /blog/new-title. Pages canonicalising to /blog/old-title should name /blog/new-title.

Example 3: Trailing slash only (not reported)

Canonical https://example.com/shoes, which redirects to https://example.com/shoes/. Both are the same stored address, so nothing is reported.

How PixyScan detects this

This is decided after the crawl, by joining each page's canonical to what the crawl found at that address.

  1. Reads the canonical of every page that answered 2xx, resolves it and puts it in the stored address form.
  2. Looks the target up among the pages this scan fetched.
  3. Reports the page when the target redirects, by either record that can say so: the target's own status is 3xx, or a stored redirect chain starts at the target and ends at a different address. The destination is included when a chain records it.

What is never judged

  • A self-referencing canonical. A page naming itself — in any spelling: a trailing slash, a fragment, a tracking parameter — is the correct and common case and is not looked up.
  • A canonical whose target this crawl did not fetch. Another domain, an address no link led to, a sitemap entry the scan's URL allowance would not stretch to, or one your exclusion rules dropped. PixyScan has no measured answer for it and does not guess. Every finding records how many canonicals were judged and how many pointed at addresses the crawl did not fetch, so a result can always be read against how much of the site it covers.
  • A canonical that is not an address. Prose in the href, or a non-http scheme. That is the malformed-canonical finding (#63).
  • A page that did not itself answer 2xx. An error page or a redirect serves no document a search engine would read a canonical from.

How the four canonical-target checks fit together

Each page's canonical is judged once, in this order:

  1. The target answered 4xx or 5xx → Canonical points at an error page (#159), and nothing else.

  2. The target redirects → Canonical points at a redirect (#160), and nothing else. In particular it is never also reported as a canonical chain: when the "chain" is a redirect, the redirect is the finding.

  3. The target answered 2xx:

    • it carries noindex → Canonical points at a noindex page (#161);
    • its own canonical names another address → Canonical target canonicalises to another page (#162).

    These two can both be reported for one page, because they are different faults on the target with different fixes.

What we store

Storage Level

Page Level — reported against the page that declares the canonical.


Database Table / Prisma Model

Read from PageSeoBasicsData (both pages), Url (the target's crawl state and status) and RedirectChain (where a redirect lands). The finding is stored on audit_issues.details.


Fields Used

Field Type Description
page_seo_basics_data.canonical_url String? The canonical declared on the reported page, and on its target
page_seo_basics_data.indexability_noindex_absent Boolean? false when the target carries noindex in a robots meta tag or an X-Robots-Tag header; null is never read as noindex
urls.url String The stored address the canonical is matched to, after resolving and normalising it
urls.status Enum Only completed or failed rows count as fetched; pending, skipped and excluded are not judged
urls.status_code Int? What the target answered. Status 0 (our own failed request) is not an answer and is not judged
redirect_chains.original_url / final_url String Where a target that redirects lands, for redirects recorded under the address that was requested

Stored Fields on the finding

Field Type Description
message / recommendation String What was found at the target, and what to change
url String The page being reported
canonicalUrl String The canonical as the page declared it
targetUrl / targetUrlId String The page the canonical resolved to
targetStatusCode Int What the target answered
redirectsTo String? Where the redirect lands, when a stored chain records it
scope String Always targets-fetched-in-this-scan
basis String A sentence stating how many canonicals were judged and how many were not
pagesWithCanonical, selfCanonical, unresolvable, targetsNotFetched, targetsJudged Int The coverage counts behind basis

Detection Dependencies

  • HTML Document — the canonical tag on both pages, and the target's robots meta tag
  • HTTP Response — the target's status code and X-Robots-Tag header
  • Redirects — the stored redirect chain, when the target redirects

Further reading