Skip to content
Issue docs

HTML canonical and HTTP Link-header canonical disagree

Importantlink_header_canonical_conflictIssue 202

What is this issue?

The page declares its canonical twice, in two places, with two different URLs:

  • in the HTML — <link rel="canonical" href="https://example.com/a">
  • in the HTTP response — Link: <https://example.com/b>; rel="canonical"

The HTTP form is valid and is how canonicals are usually set for PDFs and other non-HTML files; on an HTML page it is typically added by a CDN, a server rule or a framework. When both exist and disagree, the page is sending conflicting canonical signals.

Why it matters

  • Conflicting canonicals are ignored. Google treats the two as competing hints and may disregard both, choosing a canonical itself — so the consolidation you set up stops working.

  • Two owners, two places. The HTML canonical belongs to the template; the header belongs to the server or CDN. Fixing one without knowing about the other leaves the conflict in place, and View Source shows only half of it.

  • Header rules leak. A Link rule written for one path or file type often matches more than intended — every page under a prefix declaring the same canonical.

How to fix it

  1. Find where the header comes from. Check the response in your browser's developer tools (Network → the document request → Response Headers) or with curl -I. Look for link: with rel="canonical".

  2. Keep one declaration. For an HTML page the template's canonical is usually the right one; remove rel="canonical" from the server or CDN Link rule, leaving any preload hints it also carries.

    # Before
    link: </app.js>; rel=preload; as=script, <https://example.com/b>; rel="canonical"
    
    # After
    link: </app.js>; rel=preload; as=script
  3. If you need both, make them name exactly the same absolute URL.

Examples

Example 1: The two agree

link: <https://example.com/a>; rel="canonical"
<link rel="canonical" href="/a">

Passes. Both resolve to https://example.com/a.

Example 2: Header carries only preload hints

link: </app.js>; rel=preload; as=script, </font.woff2>; rel=preload; as=font
<link rel="canonical" href="https://example.com/a">

Passes. No canonical in the header.

Example 3: They disagree

link: <https://example.com/b>; rel="canonical"
<link rel="canonical" href="https://example.com/a">

Fails.

Example 4: Hidden among preload hints

link: </bundle.js?chunks=1,2,3>; rel=preload; as=script, <https://example.com/b>; rel=canonical
<link rel="canonical" href="https://example.com/a">

Fails. The comma inside the first URL does not split the header; the canonical is still found.

How PixyScan detects this

  1. The Link header is parsed per RFC 8288. It is a comma-separated list of link-values; commas inside <…> and inside quoted parameters are not separators. A link-value declares a canonical when its rel parameter — quoted or bare, case-insensitive, possibly several space-separated tokens — contains canonical. A header sent more than once is read as one list.

  2. The HTML side is the first canonical in the head. A canonical outside the head is ignored by search engines and is reported separately as "Canonical tag outside the head".

  3. Both are resolved against the page's served URL and compared with only the fragment removed. A trailing slash or a query parameter is a different URL.

  4. Reported when both exist and differ. The live header is read, so a long header full of preload hints never hides the canonical. The stored copy of the link header is also kept up to 8 KB (other headers are capped at 1 KB) for the same reason.

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageSeoBasicsData (page_seo_basics_data). The finding itself is stored on audit_issues.details.


Fields Used

Field Type Description
canonical_url String? The first <link rel="canonical"> href in the document, exactly as written
has_duplicate_elements Boolean? True when the page declares a canonical (or title, description, H1) more than once
duplicate_elements Json? One entry per repeated element: { element: 'canonical', value, count }

PageResponseHeader (page_response_headers) keeps the response headers:

Field Type Description
headers Json Lower-cased header name to value; the link value is kept up to 8 KB, every other header up to 1 KB
truncated_values Int How many values were cut at their cap

Stored Fields on the finding

Field Type Description
htmlCanonicalUrl String The head canonical, resolved
headerCanonicalUrl String The Link-header canonical, resolved

Detection Dependencies

  • HTML Document — the canonical link and its position
  • HTTP Response Headers — the Link header

Further reading