Near-duplicate body content on another page
What is this issue?
This check reports a page whose body text is near-identical to another page's in the same crawl.
Not byte-identical — near-identical. Almost no real duplicate is an exact copy: a printer-friendly view carries a different footer, a faceted listing reorders two filters, a session-id copy differs by one link. PixyScan measures how much of the phrasing two pages share, and reports a pair once they share the great majority of it — around 95%, though the measurement is an estimate rather than a hard line. See "Where the line actually falls" in the examples for what that means in practice.
Campaign parameters (utm_*, gclid, fbclid and the like) are stripped before a
page is recorded at all, so a URL that differs only by one of those is never two pages
to PixyScan and cannot be reported here.
Body text means the page's own prose. The navigation, header and footer are removed before anything is measured, along with script and style blocks and anything the markup declares hidden. A site-wide template is the same words on every page; measuring it would make every page on every site a duplicate of every other.
For a page to pass this check:
- No other page in the crawl says substantially the same thing, or
- The page already declares which version should rank, with
rel="canonical"pointing at it, or - The page is noindexed, so it is not competing to be indexed at all.
Example that fails: /blue-widget and /blue-widget?ref=newsletter both
serve the same 800-word product description. Two addresses, one page.
Example that passes: /blog/page/2 and /blog/page/3 both list ten posts
inside the same template. The template is identical and the posts are not, so the
pages are not duplicates — and Google does not treat paginated pages as duplicates
either.
Why it matters
Only one of the copies gets indexed
When Google finds several pages saying the same thing it groups them and picks one — its own choice, not necessarily yours — to represent the group. The rest are reported in Search Console as Duplicate without user-selected canonical and do not appear in results. If the version Google picks is not the one you promote, the address in your ads, your emails and your internal links is the one that is not there.
The group's authority is split
Links, shares and internal link equity land on whichever address a person happened to copy. Four addresses serving one page divide one page's authority four ways, so a page that could compete ranks like a quarter of itself. Consolidating them is usually the single largest ranking change available on a duplicated section.
Crawl budget is spent re-reading what it already has
Every duplicate is a URL a crawler fetches, parses and discards. On a faceted catalogue that can be thousands of URLs, and the pages that are unique wait behind them.
AI answers cite one URL, not your preferred one
An engine building an answer from your content cites the address it indexed. With no canonical you are not choosing which address that is.
Effect on your SEO health score
near_duplicate_body is graded IMPORTANT. The page is served, crawlable and
perfectly indexable — nothing is mechanically broken, which is what CRITICAL
means — but it competes worse than it should, and measurably so. That is the same
grade as duplicate titles (#66) and duplicate meta descriptions (#67), which is
deliberate: a duplicated body costs at least as much as a duplicated title.
Resolving the duplicates removes the deduction from every page in the group at
once.
How to fix it
1. Decide which version should rank. Usually the shortest, cleanest address —
the one without parameters, without /print/, without a session id. Everything
below follows from that decision.
2. Point the copies at it with rel="canonical". On every duplicate, in the
<head>:
<link rel="canonical" href="https://example.com/blue-widget" />Use one absolute URL. The canonical version must point at itself as well, or the group has no agreed owner. This is the right fix when the copies must stay reachable — a print view, a filtered listing, a tracked link.
3. Redirect instead, where a copy has no reason to exist. If nothing needs the
duplicate address, 301 it to the version you kept. A redirect is stronger than a
canonical: it is an instruction rather than a hint.
4. Noindex the copies that must stay but must not rank. A <meta name="robots" content="noindex"> takes the page out of the index while leaving it
usable. Do not combine noindex with a canonical pointing at a different page —
that tells search engines to drop this page and to credit another one with it, which
are contradictory instructions. (A noindex page whose canonical points at itself is
fine and normal.)
Not on paginated listings. /page/2, /page/3 and so on are how search engines
reach the products or posts deeper in a series. Canonicalising them all to page 1, or
noindexing them, cuts that path and can drop the deeper items out of the index
entirely. If a paginated series is reported here, the fix is to give each page
content that differs — or to accept it — never to canonicalise the series away.
5. Stop generating the duplicates at source, where you can. Serve tracking parameters from the canonical URL, do not give one product a URL under every category it belongs to, and do not publish a separate print page when a print stylesheet does the job.
6. Where the pages should be different, make them different. Locations pages, service-area pages and category pages built from one template with a name swapped are duplicates in every way that matters. Either write content that is genuinely about each place, or keep one page and list the rest on it.
What not to do
- Do not add a paragraph of filler to make a copy "different". The measurement reads phrasing, and so does Google.
- Do not block the duplicates in
robots.txt. A blocked page cannot be read, so its canonical cannot be read either, and the duplicate stays in the index without a title.
Examples
Example 1 — a tracked copy of a product page
Scenario. A newsletter links to the product page with a campaign parameter the site does not strip, so the same 800 words are served at two addresses.
Fails. Both addresses serve the identical description and neither says which one counts.
<!-- https://example.com/blue-widget -->
<head><title>Blue Widget | Example</title></head>
<!-- https://example.com/blue-widget?ref=newsletter -->
<head><title>Blue Widget | Example</title></head>Google groups the two, indexes one, and reports the other as Duplicate without user-selected canonical. Which one it keeps is not the site's decision.
Passes. The tracked address declares the plain one as canonical, and the plain one declares itself.
<!-- https://example.com/blue-widget?ref=newsletter -->
<link rel="canonical" href="https://example.com/blue-widget" />
<!-- https://example.com/blue-widget -->
<link rel="canonical" href="https://example.com/blue-widget" />Example 2 — a printer-friendly view
Scenario. Every article has a /print/ twin carrying the same body inside a
bare template.
Fails. Two indexable pages, one article. The /print/ version sometimes
outranks the real one, because it is lighter and has fewer outbound links.
<!-- https://example.com/guide/onboarding -->
<!-- https://example.com/print/guide/onboarding (same 1,200 words) -->Passes. The print view is taken out of the index and points at the article.
<!-- https://example.com/print/guide/onboarding -->
<meta name="robots" content="noindex" />Better still, delete the print page and use a print stylesheet:
<link rel="stylesheet" href="/print.css" media="print" />Example 3 — templated location pages
Scenario. Twelve "SEO services in {city}" pages generated from one template with the city name swapped.
Fails. Twelve addresses, one page. Each is 96% identical to the others, none of them ranks, and the group's links are split twelve ways.
<!-- /seo-services/leeds -->
<h1>SEO services in Leeds</h1>
<p>We provide expert SEO services in Leeds. Our Leeds team …</p>
<!-- /seo-services/bristol -->
<h1>SEO services in Bristol</h1>
<p>We provide expert SEO services in Bristol. Our Bristol team …</p>Passes — one of two ways. Either each page is genuinely about its city — local case studies, local pricing, the actual team, local coverage — so the pages stop being copies of each other. Or keep one services page and list the twelve locations on it.
What is not reported
A page that already declares its canonical. Fixed is fixed; the check does not keep reporting a duplicate the site has already resolved. Neither is a noindexed page, a page that did not return 200, or one already reported as thin content or a soft 404.
Where the line actually falls
The check does not compare pages word by word — it compares 64-bit fingerprints, so what it reports is an estimate of how much phrasing two pages share. That estimate is close on average and noisy on any single pair. Measured across hundreds of page pairs built from an identical shared block plus their own article:
| share of the page that is identical | how often the pair is reported |
|---|---|
| 85% | rare |
| 90% | roughly one pair in eight |
| 95% | about half |
| 98% and above | almost always |
Two pages that really are the same document are reported every time, because an identical body produces an identical fingerprint. The softness is confined to the middle of that table.
Read it as: pages sharing most of a template are usually left alone, pages that are nearly the same document are usually caught, and in between it is a matter of degree. If a section of your site is built from one template with a small amount of unique copy per page, expect some of those pages to be reported — and treat that as a signal worth looking at rather than a mistake, because that is genuinely what thin templated content looks like to a search engine.
Pagination is not exempted by a rule. /blog/page/2 and /blog/page/3 list
different posts, so their body text differs and they are normally well clear of the
threshold — but a paginated series whose pages really do carry near-identical text
will be reported, and that is the correct answer.
How PixyScan detects this
This is one of the few checks that cannot be answered from a single page, so it happens in two parts: the crawler measures each page as it goes, and a pass after the crawl compares the measurements.
During the crawl — fingerprinting each page
Takes the page's own prose. The same text the thin-content check counts the words of:
script,style,noscriptandtemplateblocks removed, anything the markup declares hidden removed, and the page's own navigation, masthead and footer removed. What is left is what the page is about rather than what its template is about."The page's own" is the important part. A
headerorfooterinside an article, a list item or a card is that card's own heading — a product name, a result title — and is kept. Only a header or footer belonging to the page as a whole is treated as furniture. Removing every one of them regardless would delete exactly the text that tells two listing pages apart.Breaks it into four-word phrases. Every consecutive run of four words is taken as one phrase, overlapping — "the quick brown fox", "quick brown fox jumps", and so on. Phrases rather than a word list is what stops the check firing on every page of a shop: every product page draws on the same vocabulary, and only pages with the same sentences share the same phrases.
Reduces the phrases to a 64-bit fingerprint. Unlike an ordinary checksum, this fingerprint changes a little when the page changes a little, so two fingerprints can be compared for closeness rather than only for equality. A phrase used often pulls the fingerprint harder than one used once.
Records nothing when there is too little to measure. A page with fewer than 50 words of prose gets no fingerprint at all. Below that there is not enough text for the measurement to be stable, and a page that short is the thin-content check's finding.
Chinese and Japanese are written without spaces, so for those scripts each character is counted as its own unit and a "four-word phrase" becomes a four-character one — the standard way these languages are compared. Without that, a whole clause would count as one word and no CJK page would ever be fingerprinted.
After the crawl — comparing them
Compares only pages that were fingerprinted. A page with no fingerprint is never given a verdict, in either direction — it is never treated as "measured, and unique". Every finding carries the number of pages that were compared and the number that could not be fingerprinted, so a verdict can always be read against how much of the site it actually speaks for.
Finds the pairs that agree on at least 90% of their fingerprint bits. That works out at two pages sharing roughly 95% of their phrasing. Rather than comparing every page with every other, PixyScan indexes the fingerprints in seven slices and compares only pages that share a whole slice — which cannot miss a pair at this threshold, and turns an impossible amount of work on a large site into a manageable amount.
The fingerprint estimates similarity from 64 bits rather than computing it, so the boundary is a soft edge rather than a hard line. Measured: a pair sharing 90% of its phrasing is reported roughly one time in eight, a pair sharing 95% about half the time, and a pair sharing 98% or more almost always. Two pages that really are the same document have the same fingerprint and are reported every time. The examples page shows the whole curve.
On a very large scan the comparison stops at a fixed ceiling — around 26,000 fingerprinted pages' worth of work — so that it cannot block the service. When it does, every finding it wrote says so, and there may be further near-duplicates it did not reach.
Reports one finding per affected page. Two duplicate pages are one editorial fault and two pages that each need a decision, so each gets a row listing the others and the size of the group. The count therefore means "pages affected", exactly as it does for duplicate titles.
Says "exact duplicate" when that is what it found. When another page carries exactly the same fingerprint — all 64 bits agreeing, similarity 1.0 — the finding calls the page an exact duplicate rather than near-identical, and says how many of the group are exact copies and how many are only near. Calling a copy "near-identical, 100% similar" sent readers looking for a difference that is not there.
When PixyScan stays quiet
A page in a duplicate group is not reported when:
- It already declares a canonical pointing elsewhere. The site has done the thing this check asks for. Reporting it would make the finding impossible to clear and would light up precisely the sites that handle duplication correctly.
- It is noindexed. It is not in the index, so it cannot be duplicate content in it — and noindexing a copy is one of the fixes recommended here.
- It did not serve. An error page or a redirect has no content to duplicate.
- It is already reported as thin content (#127) or as a soft 404 (#129). Those pages resemble their neighbours because they are empty; the fingerprint is measuring absence rather than duplication, and each of those checks is the row that names the fix. Reporting both would charge one editorial fault twice.
The first four of those take the page out of the comparison entirely: it is not reported, and it is not named as anybody else's duplicate either, because a page that is not in the index is not competing with anything.
The last one is different. A thin page or a soft 404 stays in the comparison and is still listed against its neighbours — it simply gets no row of its own. Dropping it would silence the other page: a full, indexable, genuinely duplicated page would go unreported whenever the page it duplicates happened to be the thin one.
Duplicate titles (#66) and duplicate meta descriptions (#67) do not stand this check down, and the reverse is also true. They are different faults with different repairs: rewriting a title does not stop two pages being the same page, and canonicalising a duplicate does not give a paginated series distinct titles.
What we store
Storage Level
Page Level
Database Table / Prisma Model
PageSeoBasicsData — the measurement, one row per page per scan
audit_issues.details — the finding
Fields Used
| Field | Type | Description |
|---|---|---|
| bodySimhash | String | 64-bit fingerprint of the page's body prose, as 16 hex characters |
bodySimhash has two states and they must not be confused:
| Value | Meaning |
|---|---|
"3f9c…" |
The page's body was fingerprinted and can be compared |
null |
No fingerprint for this page — a scan taken before the column existed, a non-HTML resource, a crawl that failed, or a body under 50 words. Never "measured and unique" |
The finding carries the comparison itself, so a report can be read without joining back to the page rows:
| Field | Type | Description |
|---|---|---|
| url | String | The page the text was read from |
| duplicateUrls | String[] | Addresses of the other pages in the group (up to 25 listed) |
| duplicateUrlIds | String[] | Their page ids, for joining |
| occurrences | Int | The size of the whole group, this page included — never trimmed |
| listingCap | Int | How many addresses the listing shows at most |
| similarity | Float | Fingerprint agreement with the closest match, 0 to 1 |
| exact | Boolean | True when at least one other page carries this page's exact fingerprint (similarity 1.0) |
| exactMatches | Int | How many of the others are exact copies; the rest of the group are near matches |
| threshold | Float | The threshold applied (0.9) |
| scope | String | pages-in-this-scan — the basis of the claim |
Detection Dependencies
- HTML Document (the response body on the plain HTTP engine; the document after scripts have run on the JavaScript engine)
- HTTP Response (the status code — a page that did not serve is not compared)
- The page's canonical tag and noindex signals, both on the same
PageSeoBasicsDatarow - The scan's existing
thin_content(#127) andsoft_404_page(#129) findings, which this check stands down for
Note
The fingerprint is stored on the page SEO basics row rather than on the
readability row, for the same reason wordCount is: duplicate content is a
page-basics question and the readability group can be switched off
independently. Both are taken from the same prose, so "this page is thin" and
"this page is a duplicate" are always statements about the same text.
The column is deliberately not indexed. Nothing looks a fingerprint up by value — the post-crawl pass reads the scan's column once and builds its own index in memory — so a database index would cost every page write and serve no read.