Sitemap lists URLs blocked by robots.txt
What is this issue?
Your XML sitemap and your robots.txt disagree. The sitemap lists URLs and asks search engines to index them; robots.txt forbids search engines to fetch those same URLs.
For a site to pass this check:
- No URL in the XML sitemap is disallowed by robots.txt for Googlebot (or for
*when the file has no Googlebot section).
Example: a site blocks /staging/ in robots.txt, but its sitemap generator walks the
whole content tree and lists 40 staging pages. Search Console reports each one as
"Submitted URL blocked by robots.txt".
Why it matters
Search Console reports it as an error. Google flags every sitemap URL it is not allowed to crawl, and a sitemap full of contradictions is one it trusts less.
The pages are not crawled. If the pages are meant to rank, robots.txt is keeping them out. If they are not, the sitemap is advertising pages you are hiding.
Blocked URLs can be indexed without content. A URL Google knows about but may not fetch can still appear in results as a bare address with no description.
Why IMPORTANT, a grade above the other sitemap checks. A sitemap URL that redirects or 404s is wasted effort. Here the site gives search engines two direct, opposite instructions about the same page, and one of them is wrong.
How to fix it
Open the finding. Each listed URL comes with the Disallow rule that blocks it and the
sitemap file that lists it. Decide, section by section, which of the two files is right.
The pages should not be in search: take them out of the sitemap. Usually that means excluding the path in your sitemap generator or CMS settings rather than editing the XML by hand.
The pages should be in search: remove or narrow the robots.txt rule.
# Before: also blocks /blog-archive/, which is in the sitemap Disallow: /blog # After Disallow: /blog/drafts/Regenerate and resubmit the sitemap in Search Console so the errors clear.
Prefer
noindexto robots.txt for pages that must stay out of the index. Search engines can only see anoindexon pages they are allowed to fetch.
Examples
Example 1: a blocked section in the sitemap
robots.txt:
User-agent: *
Disallow: /internal/Problematic sitemap:
<url><loc>https://example.com/internal/team-handbook</loc></url>Corrected: remove the entry from the sitemap, or remove the Disallow if the page
is meant to be public.
Example 2: Googlebot has its own rules
User-agent: *
Allow: /
User-agent: Googlebot
Disallow: /beta/A sitemap entry for https://example.com/beta/new-feature is reported. Google follows
its own section, even though * allows everything.
Example 3: what passes
User-agent: *
Disallow: /admin/
Allow: /admin/help<url><loc>https://example.com/</loc></url>
<url><loc>https://example.com/admin/help</loc></url>/admin/help is carved back out by the longer Allow rule, so it is crawlable.
How PixyScan detects this
robots.txt and the XML sitemap are both read during the site-level part of the scan. Sitemap indexes are followed to every child sitemap.
The section a search engine obeys is selected:
Googlebotif the file has one, otherwise*. This does not depend on the scan's "Respect robots.txt" setting.Every
<loc>is tested against that section: path and query string,*wildcards,$end anchors, longest rule wins,Allowwins a tie.One site-level finding is raised if any entry is blocked. It lists up to 100 entries, each with the rule that blocks it and the sitemap file that listed it, and gives the total.
What is deliberately not reported:
- URLs your scan's exclude or include patterns keep out of the report. You asked for those not to be reported on.
- The same URL twice. A URL listed by two sitemap files counts once.
No double reporting with the other sitemap checks. A blocked entry is reported here only. The checks for sitemap URLs that do not return 200, are noindexed, or are not canonical skip it, even when a scan that ignores robots.txt went on to fetch it. The fix is the same edit to the same line, so it is reported once, under the higher grade.
Blocked sitemap URLs are not added to the scan's page list. When the scan respects robots.txt, PixyScan does not fetch them either.
What we store
Storage Level
Site Level: the finding is about the sitemap and robots.txt together, not about any one page. It is raised once per scan with the conflicting entries listed on it.
Database Table / Prisma Model
AuditIssue (url_id null)
The blocked entries are not stored as page rows: robots.txt forbids them, so they are kept off the scan's page list. The finding is the record of them.
Finding Details
| Key | Type | Description |
|---|---|---|
blockedCount |
Number | Distinct sitemap entries robots.txt blocks (true total) |
urls |
Array | Up to 100 { url, rule, userAgentGroup, sitemapFileUrl } |
truncated |
Boolean | True when urls is a sample |
userAgent |
String | googlebot: the section applied (falls back to *) |
url is the <loc> exactly as the sitemap declares it.
Detection Dependencies
- XML sitemap: every
<loc>, with the sitemap file that listed it - robots.txt: fetched once per scan, per
User-agentsection