Skip to content
Issue docs

Canonical points at a noindex page

Importantcanonical_to_noindex_pageIssue 161

What is this issue?

This issue reports a page whose <link rel="canonical"> names a page that carries a noindex directive — in a robots meta tag or in an X-Robots-Tag header.

The canonical says "index that page instead of me". The target says "do not index me". Between them, neither may be indexed.

For a page to pass this check:

  • The page its canonical names may be indexed.

Example: /shoes?sort=price declares a canonical to /shoes, and /shoes carries <meta name="robots" content="noindex"> left over from a staging build.

Why it matters

  • Pages can disappear from search entirely. The duplicate defers to the target; the target asks to be left out. Google resolves the conflict as it sees fit, and a common outcome is that neither version ranks.
  • It is usually an accident. A noindex copied from staging, or a section noindexed without checking what canonicalises into it.
  • It affects every page pointing at the target, so one stray noindex on a category page can take a whole set of filtered views with it.

How to fix it

Decide which page should rank.

  1. If the target is meant to rank, remove its noindex — check both the robots meta tag and the X-Robots-Tag response header.
  2. If the target is deliberately noindexed, point the canonical somewhere that is meant to rank — usually the page itself.
  3. Check how the noindex got there. Staging settings, a CMS "hide from search" switch, or a server rule adding the header to a whole path.

Examples

Example 1: Category page noindexed by mistake (reported)

<!-- https://example.com/shoes?sort=price -->
<link rel="canonical" href="https://example.com/shoes">
<!-- https://example.com/shoes -->
<meta name="robots" content="noindex, follow">

Corrected: remove the noindex from /shoes.

Example 2: Noindex sent as a header (reported)

https://example.com/guides/ answers with X-Robots-Tag: noindex, and its paginated pages canonicalise to it. The header counts exactly like the meta tag.

Example 3: Target indexable (not reported)

The target answers 200 with no noindex: the canonical can do its job.

How PixyScan detects this

  1. Reads the canonical of every page that answered 2xx, resolves it and puts it in the stored address form.
  2. Looks the target up among the pages this scan fetched.
  3. Reports the page when the target answered 2xx and carries noindex, in either its robots meta tag or its X-Robots-Tag header. A target whose noindex signal was never measured is not treated as noindexed.

What is never judged

  • A self-referencing canonical. A page naming itself — in any spelling: a trailing slash, a fragment, a tracking parameter — is the correct and common case and is not looked up.
  • A canonical whose target this crawl did not fetch. Another domain, an address no link led to, a sitemap entry the scan's URL allowance would not stretch to, or one your exclusion rules dropped. PixyScan has no measured answer for it and does not guess. Every finding records how many canonicals were judged and how many pointed at addresses the crawl did not fetch, so a result can always be read against how much of the site it covers.
  • A canonical that is not an address. Prose in the href, or a non-http scheme. That is the malformed-canonical finding (#63).
  • A page that did not itself answer 2xx. An error page or a redirect serves no document a search engine would read a canonical from.

How the four canonical-target checks fit together

Each page's canonical is judged once, in this order:

  1. The target answered 4xx or 5xx → Canonical points at an error page (#159), and nothing else.

  2. The target redirects → Canonical points at a redirect (#160), and nothing else. In particular it is never also reported as a canonical chain: when the "chain" is a redirect, the redirect is the finding.

  3. The target answered 2xx:

    • it carries noindex → Canonical points at a noindex page (#161);
    • its own canonical names another address → Canonical target canonicalises to another page (#162).

    These two can both be reported for one page, because they are different faults on the target with different fixes.

What we store

Storage Level

Page Level — reported against the page that declares the canonical.


Database Table / Prisma Model

Read from PageSeoBasicsData (both pages), Url (the target's crawl state and status) and RedirectChain (where a redirect lands). The finding is stored on audit_issues.details.


Fields Used

Field Type Description
page_seo_basics_data.canonical_url String? The canonical declared on the reported page, and on its target
page_seo_basics_data.indexability_noindex_absent Boolean? false when the target carries noindex in a robots meta tag or an X-Robots-Tag header; null is never read as noindex
urls.url String The stored address the canonical is matched to, after resolving and normalising it
urls.status Enum Only completed or failed rows count as fetched; pending, skipped and excluded are not judged
urls.status_code Int? What the target answered. Status 0 (our own failed request) is not an answer and is not judged
redirect_chains.original_url / final_url String Where a target that redirects lands, for redirects recorded under the address that was requested

Stored Fields on the finding

Field Type Description
message / recommendation String What was found at the target, and what to change
url String The page being reported
canonicalUrl String The canonical as the page declared it
targetUrl / targetUrlId String The page the canonical resolved to
targetStatusCode Int What the target answered
scope String Always targets-fetched-in-this-scan
basis String A sentence stating how many canonicals were judged and how many were not
pagesWithCanonical, selfCanonical, unresolvable, targetsNotFetched, targetsJudged Int The coverage counts behind basis

Detection Dependencies

  • HTML Document — the canonical tag on both pages, and the target's robots meta tag
  • HTTP Response — the target's status code and X-Robots-Tag header
  • Redirects — the stored redirect chain, when the target redirects

Further reading