Skip to content
Issue docs

Canonical URL written as a relative path

Suggestionrelative_canonical_urlIssue 158

What is this issue?

This is advice, not a defect. The page's canonical URL is written without a scheme:

  • root-relative — <link rel="canonical" href="/shoes">
  • path-relative — href="shoes" or href="../shoes"
  • protocol-relative — href="//example.com/shoes"

Protocol-relative URLs are treated as relative here. They name the host but not the scheme, so the http:// copy of a page declares the http:// URL canonical — which is half of what an absolute canonical is there to prevent.

Google resolves relative canonicals correctly, which is why this deducts nothing. It recommends absolute URLs because a relative one depends on where the page is served.

Why it matters

  • A relative canonical follows the page around. The same HTML served from a staging host, a preview deployment, a mirror or a scraper's copy declares that host canonical. An absolute canonical keeps pointing at the real site.

  • It depends on the base URL. A <base href> tag or a proxy that rewrites the request path changes what a relative canonical resolves to.

  • Protocol-relative leaves HTTP open. If the HTTP version of a page is ever served, it canonicalises to itself rather than to HTTPS.

Effect on the health score

None. This is a suggestion: it is shown so you can harden the canonical, and it never deducts.

How to fix it

Write the canonical as a full URL, including https:// and the host.

<!-- Before -->
<link rel="canonical" href="/shoes/red" />
<link rel="canonical" href="//example.com/shoes/red" />

<!-- After -->
<link rel="canonical" href="https://example.com/shoes/red" />

Most frameworks and SEO plugins have a "site URL" setting used to build absolute canonicals; set it to the production origin rather than deriving it from the request.

Examples

Example 1: Absolute

<link rel="canonical" href="https://example.com/shoes/red" />

Passes.

Example 2: Root-relative

<link rel="canonical" href="/shoes/red" />

Reported. Resolves to https://example.com/shoes/red on the live site, and to the staging host on a staging copy.

Example 3: Protocol-relative

<link rel="canonical" href="//example.com/shoes/red" />

Reported. On http://example.com/shoes/red this resolves to the HTTP URL.

Example 4: Absolute HTTP

<link rel="canonical" href="http://example.com/shoes/red" />

Not reported here. It is absolute; the HTTP target is reported as "Canonical URL pointing at HTTP instead of HTTPS".

How PixyScan detects this

  1. The page's canonical is read. The first <link rel="canonical"> (rel matched as a case-insensitive token), with its href as written.

  2. Absolute means "has a scheme". An href starting with a scheme and a colon — https:, http:, in any case — is absolute. Everything else is reported, including //host/path.

  3. The resolved form is shown. The finding carries the href as written and the absolute URL it resolves to against the page, which is what to write instead, and flags protocol-relative values separately.

  4. Not reported: a page with no canonical (that is "Pages with no canonical tag"), or an absolute http:// canonical (that is "Canonical URL pointing at HTTP instead of HTTPS").

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageSeoBasicsData (page_seo_basics_data). The finding itself is stored on audit_issues.details.


Fields Used

Field Type Description
canonical_url String? The first <link rel="canonical"> href in the document, exactly as written
has_duplicate_elements Boolean? True when the page declares a canonical (or title, description, H1) more than once
duplicate_elements Json? One entry per repeated element: { element: 'canonical', value, count }

Stored Fields on the finding

Field Type Description
canonicalUrl String The href as written
resolvedCanonicalUrl String? The absolute URL it resolves to against the page
protocolRelative Boolean True for //host/path

Detection Dependencies

  • HTML Document — the canonical link's href

Further reading