Skip to content
Issue docs

URL listed in more than one sitemap file

Suggestionsitemap_url_in_multiple_filesIssue 150

What is this issue?

When a site publishes several sitemap files under an index, each page should appear in exactly one of them. This check reports pages that appear in two or more.

For a site to pass this check:

  • No URL is listed by more than one sitemap file.

This is advice, not a defect. Search engines deduplicate the <loc> set before doing anything with it, so nothing about crawling or ranking changes. What a duplicated entry costs is the site's own ability to reason about its sitemaps.

Example: a blog publishes sitemap-posts.xml and sitemap-categories.xml. A post that belongs to two categories is emitted by both generators, so it appears twice — with a different <lastmod> in each, because the two generators compute it differently.

Why it matters

Two files can disagree about the same page. <lastmod>, <priority> and <changefreq> are per-entry, so a page listed twice can carry two different answers. Search engines pick one, and which one is not something you control.

Per-file coverage stops adding up. Search Console reports discovered and indexed counts per sitemap file. When pages appear in several files those numbers overlap, and "we submitted 12,000 URLs" is no longer a number you can compare with anything.

It is usually a symptom. Duplicated entries almost always mean two generators are running over overlapping sets — which tends to mean neither of them owns the answer to "which pages should be in the sitemap".

Why this is advice rather than a defect. There is no measurable search cost to point at. Google deduplicates and moves on. A check with no fault behind it is a suggestion: it carries no severity, and however widely it is raised it does not move the health score.

How to fix it

  1. Read the list on the finding. Each entry names the duplicated URL and the sitemap files that list it.

  2. Give each page one owner. Decide which file a page belongs in — usually by content type — and have every other generator exclude it.

  3. Partition rather than filter. The reliable arrangement is one query that produces the full URL set, split into files by a deterministic rule (content type, or a hash, or simply sequential batches). Several independent queries that happen to overlap is how this arises.

  4. Make sure the <lastmod> comes from one place. Where a page legitimately has to be in two files, at least ensure both generators read the same modification timestamp — two different answers is the part that actually costs you something.

  5. Note that two addresses serving the same file are not this finding. If /sitemap.xml and /sitemap_index.xml return the same document, that is one file reachable two ways and PixyScan does not report it.

Examples

Example 1 — two generators, overlapping sets

Problematic: a post in two categories is emitted by both generators, with two different <lastmod> values.

<!-- sitemap-posts.xml -->
<url>
  <loc>https://example.com/blog/spring-guide</loc>
  <lastmod>2026-03-01</lastmod>
</url>

<!-- sitemap-categories.xml -->
<url>
  <loc>https://example.com/blog/spring-guide</loc>
  <lastmod>2026-01-14</lastmod>
</url>

Corrected: posts live in sitemap-posts.xml; the category file lists category pages only.

<!-- sitemap-categories.xml -->
<url><loc>https://example.com/blog/category/guides</loc></url>

Example 2 — what is NOT reported

/sitemap.xml and /sitemap_index.xml return byte-identical documents. That is one file with two addresses, and none of its URLs is reported.


Example 3 — the query string is part of the address

<!-- sitemap-1.xml -->
<url><loc>https://example.com/blog?page=1</loc></url>

<!-- sitemap-2.xml -->
<url><loc>https://example.com/blog?page=2</loc></url>

Two different pages in two different files. Nothing is reported. A trailing slash, though, is not a difference: /a in one file and /a/ in another is the same URL listed twice.

How PixyScan detects this

Answered from the sitemap alone, during the site-level pass.

  1. Walk every sitemap file, following the index and the conventional locations, and keep every entry alongside the file that listed it.

  2. Identify the files by their contents, not by their addresses. A great many sites answer more than one conventional path — /sitemap.xml and /sitemap_index.xml are a stock CMS default — with the same document. Counted naively, every URL on such a site would look like it was in two sitemaps. PixyScan hashes each file's body, so two addresses returning one document count as one file. A file whose body could not be hashed is treated as its own file rather than lumped in with every other unmeasured one.

  3. Compare the addresses the way the crawl does — fragment dropped, a trailing slash dropped from the path, the query kept. ?page=1 and ?page=2 are two pages; /a and /a/ are one.

  4. Report every URL claimed by two or more distinct files, listing one address per file.

What is deliberately never reported:

  • A URL repeated inside a single file. That is one file's own duplicate entry, not the cross-file ambiguity this check is about.
  • A URL whose source file was not recorded.

A limit worth knowing: PixyScan reads up to 5,000 entries from each sitemap file, so on a very large file the list of duplicates is a floor rather than a complete census.

What we store

Storage Level

Site Level — the finding is about the sitemap, not about any one page in it, so it is raised once per scan with the offending entries listed on it.


Database Table / Prisma Model

AuditIssue (url_id null), read from the sitemap walk, and the file tree stored on SiteCrawlBehaviourData.


Fields Used

Field Type Description
SiteCrawlBehaviourData.sitemapFileTree Json Every sitemap file the walk found, and which index listed it
(walk) entry.loc String The <loc> as declared
(walk) entry.sourceSitemapUrl String The file that listed it — kept per entry, before deduplication
(walk) file.contentHash String? Identifies one document served at two addresses as one file

Detection Dependencies

  • XML Sitemap — every file, including nested index files, and every entry in each
  • HTTP Response — each file's body, which is what the content hash is taken over

Further reading