Sitemap lists URLs canonicalised to a different address
What is this issue?
A sitemap should list canonical URLs — one entry per page, at the address you want that page to be known by.
This check reports listed addresses whose own rel="canonical" names a different address.
The page is telling search engines "index that one instead", while the sitemap is asking
them to index this one.
For a sitemap to pass this check:
- Every URL it lists is its own canonical.
Example: a shop's faceted navigation produces /shoes?colour=red, which correctly
canonicalises to /shoes. The sitemap generator walks the site's routes and includes every
variant it finds, so /shoes?colour=red is in the file — an address the site itself has
already said is not the one that matters.
Why it matters
Every listed variant is a fetch that ends in a discard. Search engines request the address, read the canonical, and index the other page. The entry achieved a request and nothing else.
It makes the sitemap's size meaningless. "We submitted 40,000 URLs and 6,000 are indexed" is a report a site owner cannot act on if 30,000 of those were never candidates.
Mixed signals slow consolidation. Canonicalisation is a hint, not a directive. Repeating a non-canonical address in the file whose job is to nominate pages is evidence pointing the other way, and it is evidence you control.
Why this is graded STANDARD. The canonical mismatch itself — a page whose canonical points somewhere else — is already reported and scored per page by the check that owns it. What this one adds is that the sitemap is nominating the variant anyway: real, indirect, and fixed by deleting a line. Charging it at full weight would bill one canonical decision twice.
Fixing it makes the sitemap's URL count a number that means something.
How to fix it
Read the list on the finding. Each entry names the listed address, the canonical it declares, and the sitemap file it came from.
Replace the variant with its canonical, and remove the duplicate if the canonical is already listed. There should be exactly one entry per page.
Filter the generator on the canonical, not on the route. These entries appear because the sitemap is built from "every URL that resolves" rather than "every URL we nominate". Emitting an entry only when a page's canonical is itself removes the whole class permanently.
Leave the canonical tags alone unless they are wrong. In almost every case the canonical is correct and the sitemap is the thing to change. If the variant genuinely is the page you want indexed, fix the canonical instead — but not both.
Watch for parameter and pagination variants specifically. Tracking parameters, faceted filters and
?page=2URLs are where this comes from on most sites.
Examples
Example 1 — a faceted variant in the sitemap
Problematic:
<url><loc>https://example.com/shoes?colour=red</loc></url>
<url><loc>https://example.com/shoes</loc></url><!-- https://example.com/shoes?colour=red -->
<link rel="canonical" href="https://example.com/shoes" />Corrected: one entry, for the page that is nominated.
<url><loc>https://example.com/shoes</loc></url>Example 2 — a paginated listing
Problematic: every page of a listing canonicalises to page one, and every page of it is in the sitemap.
<url><loc>https://example.com/blog?page=2</loc></url>
<url><loc>https://example.com/blog?page=3</loc></url><link rel="canonical" href="https://example.com/blog" />Corrected: list /blog once. (If the deeper pages are meant to be indexed on their own,
the fix is to make each page self-canonical — but then the sitemap was not the problem.)
Example 3 — what is NOT reported
A canonical that differs only in ways the crawl treats as the same address.
<url><loc>https://example.com/about</loc></url><link rel="canonical" href="https://example.com/about/" />
<!-- or href="/about" — resolved against the page first -->
<!-- or href="https://example.com/about?utm_source=newsletter" -->All three are the same page, and none of them is a finding.
How PixyScan detects this
A page's canonical is only knowable once the page has been fetched, so this runs after the crawl rather than while the sitemap is being walked.
Take every URL the sitemap listed.
Keep only the ones that answered HTTP 200. A page that did not serve has no canonical worth reading, and it is a different finding.
Read the canonical the page declares. A relative
hrefis resolved against the page first —<link rel="canonical" href="/shoes">is ordinary markup and comparing it raw against an absolute address would report every page that uses one.Compare it with the address the page is served from, through the same normalisation the crawl files pages under: the fragment is dropped, a trailing slash is dropped from the path, campaign parameters are dropped and the remaining query is sorted. So a canonical that differs only by a trailing slash, a fragment or a
utm_sourceis the same page and is not reported.Report every listed address whose canonical names something else, together with the canonical it names and the sitemap file it came from.
What is deliberately never reported:
- A page with no canonical tag at all. That is a different check.
- A page whose canonical was never measured — no page analysis recorded, or the address was never fetched.
A URL already reported as not answering 200 is not also reported here.
What we store
Storage Level
Site Level — the finding is about the sitemap, not about any one page in it, so it is raised once per scan with the offending entries listed on it.
Database Table / Prisma Model
AuditIssue (url_id null), read from Url joined against PageSeoBasicsData.
Fields Used
| Field | Type | Description |
|---|---|---|
| Url.source | Enum | SITEMAP or BOTH — the sitemap listed this address |
| Url.sitemapLoc | String? | The <loc> verbatim, which is what the finding names |
| Url.url | String | The address the page is served from |
| Url.statusCode | Int? | Only pages that answered 200 are judged |
| PageSeoBasicsData.canonicalUrl | String? | The canonical the page declares |
Detection Dependencies
- XML Sitemap — the
<loc>list - HTML Document — the
link rel="canonical"tag in the page head