Skip to content
Issue docs

Paginated page canonicalised to the first page

Standardpaginated_canonical_to_first_pageIssue 163

What is this issue?

The page is page 2 or later of a paginated series — a category listing, a blog archive, search results — and its canonical points at page 1:

<!-- https://example.com/shoes?page=3 -->
<link rel="canonical" href="https://example.com/shoes" />

Page 3 is not a duplicate of page 1: it lists different items. Canonicalising it to page 1 tells search engines to treat it as a copy.

Why it matters

  • Items on later pages lose their path into the index. Search engines fold a canonicalised page into its target and crawl it less. The products or articles that only appear on page 3 are linked from a page that is being treated as a duplicate.

  • Google says not to do it. Google's pagination guidance: don't use the first page of a paginated sequence as the canonical page; give each page its own canonical URL.

  • It is easy to ship by accident. A template that builds the canonical from the path and drops the query string produces exactly this on every paginated URL.

How to fix it

Give every page of the series a self-referencing canonical, including its page number.

<!-- https://example.com/shoes?page=3 -->

<!-- Before -->
<link rel="canonical" href="https://example.com/shoes" />

<!-- After -->
<link rel="canonical" href="https://example.com/shoes?page=3" />

In most templates the cause is a canonical built from the path alone. Keep the page parameter when building it, while still dropping sort orders, tracking parameters and session IDs.

If you publish a view-all page that lists every item, canonicalising the paginated pages to the view-all URL is an accepted alternative.

Examples

Example 1: Self-referencing pages

<!-- https://example.com/shoes?page=3 -->
<link rel="canonical" href="https://example.com/shoes?page=3" />

Passes.

Example 2: Page 3 canonicalised to page 1

<!-- https://example.com/shoes?page=3 -->
<link rel="canonical" href="https://example.com/shoes" />

Fails. ?page=1 as the canonical fails the same way.

Example 3: A blog archive

<!-- https://example.com/blog/page/4/ -->
<link rel="canonical" href="https://example.com/blog/" />

Fails.

Example 4: Paginated without a page number in the URL

<!-- https://example.com/shoes/more -->
<link rel="prev" href="https://example.com/shoes" />
<link rel="canonical" href="https://example.com/shoes" />

Fails. The prev link says this is a later page, and the canonical points at the first.

Example 5: A view-all canonical

<!-- https://example.com/shoes?page=3 -->
<link rel="canonical" href="https://example.com/shoes/all" />

Passes. Canonicalising to a view-all page is an accepted pattern.

Example 6: A WordPress post

<!-- https://example.com/?p=3 -->
<link rel="canonical" href="https://example.com/" />

Not reported here. ?p= is a post id, not a page number.

How PixyScan detects this

  1. Is this page 2 or later? Two kinds of evidence:

    • the URL carries a page number above 1 — a page, paged, pg, pagenum, pageno, page_no, page_num, pagenumber or page_number query parameter (name in any case), or a /page/N path ending. ?p= is deliberately not read: it is WordPress's post id far more often than a page number;
    • otherwise, the page declares <link rel="prev">, so it is not the first page.
  2. What is page 1? For a numbered URL, the same URL with the page marker removed. For rel="prev" evidence, the prev URL with its page marker removed — which is the prev URL itself when this is page 2.

  3. Does the canonical point there? The canonical is resolved against the page, its own page marker removed, and compared with page 1 — fragment dropped, other query parameters sorted, trailing slash ignored. It matches only if it carries no page number or page number 1.

  4. Not reported: a self-referencing canonical; a canonical to a different page number (page 3 to page 2); a canonical to a separate view-all URL, which Google accepts.

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageSeoBasicsData (page_seo_basics_data). The finding itself is stored on audit_issues.details.


Fields Used

Field Type Description
canonical_url String? The first <link rel="canonical"> href in the document, exactly as written
has_duplicate_elements Boolean? True when the page declares a canonical (or title, description, H1) more than once
duplicate_elements Json? One entry per repeated element: { element: 'canonical', value, count }
pagination_prev String? The <link rel="prev"> href, as written
pagination_next String? The <link rel="next"> href, as written

Stored Fields on the finding

Field Type Description
canonicalUrl String The canonical href as written
firstPageUrl String Page 1 of the series, as compared
pageNumber Int? This page's number when the URL carries one
evidence String url (page number in the URL) or rel-prev

Detection Dependencies

  • Page URL — the page number
  • HTML Document — the canonical link and <link rel="prev">

Further reading