Skip to content
Issue docs

noindex page with a canonical pointing elsewhere

Importantnoindex_with_external_canonicalIssue 155

What is this issue?

The page is noindexed — by a robots meta tag or the X-Robots-Tag header — and at the same time its canonical points at a different URL.

Those are two different instructions:

  • noindex says drop this page from search results.
  • a canonical to another URL says this page is a duplicate of that one; consolidate its signals there.

Combined, they contradict each other, and search engines have to guess which one you meant.

Example: /shoes/red?sort=price carries <meta name="robots" content="noindex"> and <link rel="canonical" href="https://example.com/shoes/red">. Both were added to keep the sorted variant out of the index — but only one of them should be.

Why it matters

  • The noindex can travel to the canonical target. Google has said that when a page is both noindexed and canonicalised elsewhere, it may treat the canonical target as noindexed too. The page you meant to strengthen can drop out of the index.

  • The signals cancel out. A canonical asks for this page's signals to be merged into the target; a noindex asks for the page to be discarded. Search engines cannot do both, so the consolidation you wanted may simply not happen.

  • It hides which rule is the real one. These are usually written by two different people or plugins. The next person to touch the page cannot tell which instruction is intended.

How to fix it

Decide what you want for the page, and keep only the instruction that says it.

  1. The page is a duplicate of another page (a sorted, filtered or tracking variant): keep the canonical, remove the noindex.

    <!-- Before -->
    <meta name="robots" content="noindex" />
    <link rel="canonical" href="https://example.com/shoes/red" />
    
    <!-- After -->
    <link rel="canonical" href="https://example.com/shoes/red" />
  2. The page should not be in search at all (a thank-you page, an internal search result): keep the noindex, and make the canonical self-referencing or remove it.

    <meta name="robots" content="noindex, follow" />
    <link rel="canonical" href="https://example.com/thank-you" />
  3. Check both places. If the noindex comes from an X-Robots-Tag header or the canonical from a Link header, the change belongs in the server or CDN configuration, not the template.

Examples

Example 1: A duplicate handled with a canonical only

<!-- https://example.com/shoes/red?sort=price -->
<link rel="canonical" href="https://example.com/shoes/red" />

Passes. One instruction: consolidate into /shoes/red.

Example 2: A hidden page with a self-canonical

<!-- https://example.com/thank-you -->
<meta name="robots" content="noindex" />
<link rel="canonical" href="/thank-you" />

Passes. The canonical names the page itself, so there is no contradiction.

Example 3: Both instructions on one page

<!-- https://example.com/shoes/red?sort=price -->
<meta name="robots" content="noindex" />
<link rel="canonical" href="https://example.com/shoes/red" />

Fails. Drop the page, and consolidate it into another — pick one.

Example 4: Delivered over HTTP

HTTP/2 200
x-robots-tag: noindex
link: <https://example.com/shoes>; rel="canonical"

Fails. The same contradiction, set in the server configuration.

How PixyScan detects this

  1. Is the page noindexed? The same verdict as "Pages that search engines cannot index" (#21): noindex or none in the robots meta tag or the X-Robots-Tag response header. The two checks can never disagree about it.

  2. Does a canonical point elsewhere? Both declaration routes are read, in this order: the HTML <link rel="canonical">, then a rel="canonical" in the HTTP Link header.

  3. "Elsewhere" is the self-canonical comparison (#73). The canonical is resolved against the page URL, and the fragment, the query string and a trailing slash are ignored on both sides. A canonical to the same page with a tracking parameter, a relative self-canonical, or one differing only by a trailing slash is not reported.

  4. One finding per page, naming the canonical target and which route declared it.

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageSeoBasicsData (page_seo_basics_data). The finding itself is stored on audit_issues.details.


Fields Used

Field Type Description
x_robots_tag String? The X-Robots-Tag response header verbatim; '' when the response was read and carried none, NULL when no response headers were read
indexability_noindex_absent Boolean? #21's verdict: false when the first robots meta tag or the header carries noindex or none
canonical_url String? The first <link rel="canonical"> href, as written

The robots meta content #26 reads is also stored, on PageHtmlHeadAudit (robots_meta_content) — the first name="robots" tag only.


Stored Fields on the finding

Field Type Description
canonicalUrl String The canonical target, resolved and normalised as compared
canonicalSource String html or link-header
robotsMeta String? The robots meta content that #21 read
xRobotsTag String? The X-Robots-Tag header

Detection Dependencies

  • HTML Document — the robots meta tag and the canonical link
  • HTTP Response Headers — X-Robots-Tag and Link

Further reading