Skip to content
Issue docs

Canonical URL not present in the XML sitemap

Importantcanonical_matches_exactlyIssue 16

What is this issue?

This issue reports a page whose canonical address is not listed in your XML sitemap.

Every indexable page has a canonical address — the one version of it you want search engines to index. The XML sitemap is your list of exactly those addresses. When a page's canonical is missing from the sitemap, the two signals you send about which pages matter disagree.

The canonical address is:

  • the URL in the page's <link rel="canonical"> tag, when it has one; or
  • the page's own address, when it declares no canonical at all. A page with no canonical tag is still its own canonical in Google's eyes, so it is held to the same rule.

For a page to pass this check:

  • Its canonical address (declared or implied) appears as a <loc> in the sitemap.

Example:

  • Canonical tag: <link rel="canonical" href="https://example.com/products/shoes">
  • Sitemap entry: <loc>https://example.com/products/shoes</loc>
  • ✅ Listed — passes

Example of a failure:

  • https://example.com/products/boots declares no canonical and answers 200
  • The sitemap lists /products/shoes but not /products/boots
  • ❌ The page's own address is its canonical, and it is missing from the sitemap

Spelling differences that do not change the page do not count as a mismatch: a trailing slash, http versus https, upper-case letters in the host name, a #fragment, tracking parameters such as utm_source, and the order of query parameters. A different host name (www. versus bare) or a different path does.

Why it matters

Google's Canonical Selection: When the canonical tag and sitemap URL don't match, Google may choose a different canonical than intended, potentially indexing the wrong version of the page.

Mixed Signals: Search engines receive conflicting information about which URL is the preferred version, leading to confusion and potential indexing issues.

Crawl Budget Waste: If search engines try to reconcile the mismatch, they may spend unnecessary resources crawling multiple versions.

Link Equity Consolidation: Inconsistent canonicals prevent proper consolidation of ranking signals to the intended URL.

Pages left out of the sitemap are discovered later. A page with no canonical tag that is also missing from the sitemap depends entirely on internal links to be found and recrawled.

SEO Health Score: Resolving this issue improves the technical SEO score by ensuring consistent canonical signals across sitemap and on-page elements.

How to fix it

  1. Audit current state - Compare the canonical tags on your pages with the URLs listed in your XML sitemap.

  2. Standardize your preferred URL format - Decide on the canonical URL format considering:

    • Protocol (always https)
    • www vs non-www (be consistent)
    • Trailing slashes (be consistent)
    • Query parameters (usually stripped)
  3. Update canonical tags - Ensure all canonical tags use the exact preferred URL format.

  4. Update XML sitemap - Ensure all URLs in the sitemap match the canonical URL format exactly.

  5. List every indexable page, including those with no canonical tag - A page without a canonical is its own canonical. Either add it to the sitemap, or, if it should not rank, noindex it. Adding a self-referencing canonical makes the intent explicit.

  6. Check for dynamic generation - If canonicals or sitemaps are generated dynamically, ensure the logic produces identical URLs.

  7. Validate with tools - Use Google Search Console's URL Inspection tool to verify Google's selected canonical matches your intended canonical.

Examples

Example 1: Canonical listed, in a different spelling (passes)

  • Sitemap: <loc>http://example.com/page/</loc>
  • Canonical: <link rel="canonical" href="https://example.com/page">

Passes. Scheme and trailing slash are spelling differences, not different pages. Use one spelling everywhere anyway; it removes doubt.


Example 2: www vs non-www (fails)

  • Sitemap: <loc>https://example.com/page</loc>
  • Canonical: <link rel="canonical" href="https://www.example.com/page">

Fails. www.example.com and example.com are different hosts.

Corrected:

  • Sitemap: <loc>https://www.example.com/page</loc>
  • Canonical: <link rel="canonical" href="https://www.example.com/page">

Example 3: A page with no canonical, missing from the sitemap (fails)

  • https://example.com/guides/returns answers 200, has no canonical tag and no noindex.
  • The sitemap does not list it.

Fails. With no tag, the page is its own canonical, and that address is not in the sitemap.

Corrected: add <loc>https://example.com/guides/returns</loc> to the sitemap, and add a self-referencing canonical while you are there:

<link rel="canonical" href="https://example.com/guides/returns">

Example 4: A noindexed page with no canonical (not reported)

  • https://example.com/cart carries <meta name="robots" content="noindex"> and is not in the sitemap.

Not reported. A page that asks not to be indexed should not be in the sitemap.

How PixyScan detects this

This check runs after the crawl, once the sitemap has been read and every page has been analysed.

  1. Collects the pages your sitemap lists. The sitemap's <loc> entries are stored as page rows marked as coming from the sitemap. Both the address as the sitemap wrote it and the normalised address the crawl filed it under count as listed.

  2. Works out each page's canonical address.

    • A page with a <link rel="canonical"> uses that URL, resolved against the page first, so a relative canonical such as /shoes is read as https://example.com/shoes.
    • A page without a canonical uses its own address — but only when it is asking to be indexed: it answered with a 2xx status and carries no noindex (neither in a robots meta tag nor in an X-Robots-Tag header). A redirect, an error page or a noindexed page should not be in a sitemap, so it is not held to this rule. A page whose status or noindex signal was never measured is not judged either.
  3. Normalises both sides the same way. Fragments, trailing slashes, tracking parameters and query-parameter order are dropped or sorted, the host is lower-cased and http is treated as https. Only a difference that names a different page counts.

  4. Reports a page whose canonical address is not among the listed pages.

  5. Says nothing when the sitemap listed no pages at all. With nothing listed, every page would fail; that is the missing-sitemap finding, raised once, not this one raised on every page.

The large-sitemap case is handled without loading the sitemap into memory: the listed pages are read in batches and the pass stops as soon as every canonical it was looking for has been found.

What we store

Storage Level

Page Level — This issue is evaluated for each individual page.


Database Table / Prisma Model

PageSeoBasicsData, read together with the page's Url row and the scan's sitemap rows in Url. The finding is stored on audit_issues.details.


Fields Used

Field Type Description
page_seo_basics_data.canonical_url String? The canonical declared on the page. When empty, the page's own address is used instead
page_seo_basics_data.canonical_matches_sitemap Boolean? Written for pages that DECLARE a canonical: true when it is listed, false when not, null when the sitemap listed nothing. Stays null for a page with no canonical
page_seo_basics_data.indexability_noindex_absent Boolean? Read for a page with no canonical: only true (no noindex in meta or header) makes it eligible
urls.url String The page's own address, used as its canonical when it declares none
urls.status_code Int? Read for a page with no canonical: only a 2xx page is eligible
urls.source / urls.sitemap_loc Enum / String? Pages the sitemap listed (SITEMAP or BOTH), as stored and as declared

Stored Fields on the finding

Field Type Description
message String What was not found in the sitemap
canonicalUrl String? The declared canonical; null for a page that declared none
impliedCanonicalUrl String Only for a page with no canonical: its own address, which was looked for
canonicalDeclared Boolean Whether the page declared a canonical tag
canonicalMatchesSitemap Boolean Always false on a finding

Detection Dependencies

  • HTML Document — the canonical tag and the robots meta tag
  • HTTP Response — the status code and the X-Robots-Tag header
  • XML Sitemap — the pages it lists

Further reading