Sitemap lists URLs that do not return HTTP 200
What is this issue?
An XML sitemap is the site telling search engines "these are my pages; please fetch them". Every address in it is a request for a crawl.
This check reports the addresses that were fetched and did not answer HTTP 200 — they 404, they 410, they 500, or they redirect somewhere else.
For a sitemap to pass this check:
- Every URL it lists serves the page it names, directly, with a 200.
A redirect counts. A 301 is the right thing to do when a page moves, but a sitemap is supposed to list the address that is the page, not an address that forwards to it — otherwise every crawl of your sitemap costs an extra request per entry and the file describes a site that no longer exists.
Example: a shop reorganises its catalogue and redirects two hundred old product URLs to their new homes. The redirects are correct. The sitemap was generated before the move and still lists the old addresses, so every entry in it is a signpost rather than a page.
Why it matters
Crawl budget is spent on nothing. Search engines fetch what a sitemap lists. Entries that 404 or redirect consume requests that could have gone to pages that exist.
Search Console reports the file as faulty. Google surfaces sitemap errors per file, and a file with dead entries reads as a file that is not maintained — which affects how much weight the rest of it is given.
It hides real changes. A sitemap that has been wrong for months is one nobody is reading, so when a page genuinely disappears there is no signal in it worth noticing.
Why this is graded STANDARD rather than higher. The pages themselves being broken is already reported, and scored, by the checks that own that fault: a page that cannot be indexed, and a broken page the site's own links point at. What this check adds is the fact that your discovery file is sending crawlers there — real, but indirect, and cheap to fix. Charging it at full weight would bill one broken page twice.
Fixing it makes the sitemap describe the site as it is now.
How to fix it
Read the list on the finding. Each entry names the address exactly as your sitemap declares it, the status it returned, and the sitemap file it came from.
For a redirect, list the destination. Replace the old address with the one it forwards to. There should be one entry per page, and it should be the address the page is actually served from.
For a 404 or 410, remove the entry. A page that is gone does not belong in a sitemap. If it should not be gone, restore it — the sitemap has told you something useful.
For a 5xx, fix the page. A server error is a fault in its own right; this check is telling you the sitemap is advertising it.
Regenerate rather than edit. A sitemap that has drifted from the site is usually one maintained by hand or by a cache that is not being invalidated. Regenerating it from the live routing table on every publish is what stops it drifting again.
Re-submit the sitemap in Search Console so the error report clears.
Examples
Example 1 — the sitemap lists a redirect
Problematic:
<url><loc>https://example.com/products/old-name</loc></url>GET /products/old-name → 301 → /products/new-nameCorrected: list the address that answers.
<url><loc>https://example.com/products/new-name</loc></url>Example 2 — the sitemap lists a page that is gone
Problematic: a seasonal landing page was deleted and returns 410, but the sitemap was not regenerated.
<url><loc>https://example.com/black-friday-2025</loc></url>Corrected: remove the entry. If the page is meant to still exist, restore it instead — the sitemap has just told you it is missing.
Example 3 — what a passing sitemap looks like
Every entry is an address that serves a page directly.
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url><loc>https://example.com/</loc></url>
<url><loc>https://example.com/products/new-name</loc></url>
<url><loc>https://example.com/about</loc></url>
</urlset>How PixyScan detects this
The sitemap knows which addresses it lists. Only the crawl knows what those addresses answer. So this check runs once, after the crawl, joining the two.
Take every URL the sitemap listed. These are the pages recorded as coming from the sitemap — both those the crawl also reached by following links, and those the sitemap was the only source for.
Read the status code the crawl received at each of them. PixyScan fetches every sitemap-listed address, including the ones no internal link leads to, so there is a status for each.
Report every listed address whose status is not 200 — any 3xx, 4xx or 5xx — naming the address exactly as the sitemap declared it, the status, and the sitemap file it came from.
What is deliberately never reported:
- A listed address the crawl never measured. If the run stopped at its page budget, was cancelled, or ran out of URL credit, some entries have no status code. "We did not look" is not "it is broken", and the count of unmeasured entries is returned rather than guessed at.
- A URL the scan's own exclude patterns or the URL quota removed. Neither is a statement about the sitemap.
A URL reported here is not also reported as noindexed or as canonicalised elsewhere: a page that did not serve has no directives and no canonical worth reading, and reporting it three times would be one fault charged three times.
What we store
Storage Level
Site Level — the finding is about the sitemap, not about any one page in it, so it is raised once per scan with the offending entries listed on it.
Database Table / Prisma Model
AuditIssue (url_id null), read from Url joined against the scan.
Fields Used
| Field | Type | Description |
|---|---|---|
| Url.source | Enum | SITEMAP or BOTH — the sitemap listed this address |
| Url.sitemapLoc | String? | The <loc> verbatim, which is what the finding names |
| Url.url | String | The normalised address the crawl filed the page under |
| Url.statusCode | Int? | What the crawl received; null means the address was not measured |
| Url.sitemapFileUrl | String? | Which sitemap file listed it |
Detection Dependencies
- XML Sitemap — the
<loc>list, walked during the site-level pass - HTTP Response — the status code recorded for each listed address
- The whole crawl — including the sitemap pass, which fetches listed pages that no internal link reaches