Skip to content
Issue docs

Sitemap lists URLs blocked by robots.txt

Importantsitemap_lists_robots_blocked_urlIssue 195

What is this issue?

Your XML sitemap and your robots.txt disagree. The sitemap lists URLs and asks search engines to index them; robots.txt forbids search engines to fetch those same URLs.

For a site to pass this check:

  • No URL in the XML sitemap is disallowed by robots.txt for Googlebot (or for * when the file has no Googlebot section).

Example: a site blocks /staging/ in robots.txt, but its sitemap generator walks the whole content tree and lists 40 staging pages. Search Console reports each one as "Submitted URL blocked by robots.txt".

Why it matters

Search Console reports it as an error. Google flags every sitemap URL it is not allowed to crawl, and a sitemap full of contradictions is one it trusts less.

The pages are not crawled. If the pages are meant to rank, robots.txt is keeping them out. If they are not, the sitemap is advertising pages you are hiding.

Blocked URLs can be indexed without content. A URL Google knows about but may not fetch can still appear in results as a bare address with no description.

Why IMPORTANT, a grade above the other sitemap checks. A sitemap URL that redirects or 404s is wasted effort. Here the site gives search engines two direct, opposite instructions about the same page, and one of them is wrong.

How to fix it

Open the finding. Each listed URL comes with the Disallow rule that blocks it and the sitemap file that lists it. Decide, section by section, which of the two files is right.

  1. The pages should not be in search: take them out of the sitemap. Usually that means excluding the path in your sitemap generator or CMS settings rather than editing the XML by hand.

  2. The pages should be in search: remove or narrow the robots.txt rule.

    # Before: also blocks /blog-archive/, which is in the sitemap
    Disallow: /blog
    
    # After
    Disallow: /blog/drafts/
  3. Regenerate and resubmit the sitemap in Search Console so the errors clear.

  4. Prefer noindex to robots.txt for pages that must stay out of the index. Search engines can only see a noindex on pages they are allowed to fetch.

Examples

Example 1: a blocked section in the sitemap

robots.txt:

User-agent: *
Disallow: /internal/

Problematic sitemap:

<url><loc>https://example.com/internal/team-handbook</loc></url>

Corrected: remove the entry from the sitemap, or remove the Disallow if the page is meant to be public.


Example 2: Googlebot has its own rules

User-agent: *
Allow: /

User-agent: Googlebot
Disallow: /beta/

A sitemap entry for https://example.com/beta/new-feature is reported. Google follows its own section, even though * allows everything.


Example 3: what passes

User-agent: *
Disallow: /admin/
Allow: /admin/help
<url><loc>https://example.com/</loc></url>
<url><loc>https://example.com/admin/help</loc></url>

/admin/help is carved back out by the longer Allow rule, so it is crawlable.

How PixyScan detects this

  1. robots.txt and the XML sitemap are both read during the site-level part of the scan. Sitemap indexes are followed to every child sitemap.

  2. The section a search engine obeys is selected: Googlebot if the file has one, otherwise *. This does not depend on the scan's "Respect robots.txt" setting.

  3. Every <loc> is tested against that section: path and query string, * wildcards, $ end anchors, longest rule wins, Allow wins a tie.

  4. One site-level finding is raised if any entry is blocked. It lists up to 100 entries, each with the rule that blocks it and the sitemap file that listed it, and gives the total.

What is deliberately not reported:

  • URLs your scan's exclude or include patterns keep out of the report. You asked for those not to be reported on.
  • The same URL twice. A URL listed by two sitemap files counts once.

No double reporting with the other sitemap checks. A blocked entry is reported here only. The checks for sitemap URLs that do not return 200, are noindexed, or are not canonical skip it, even when a scan that ignores robots.txt went on to fetch it. The fix is the same edit to the same line, so it is reported once, under the higher grade.

Blocked sitemap URLs are not added to the scan's page list. When the scan respects robots.txt, PixyScan does not fetch them either.

What we store

Storage Level

Site Level: the finding is about the sitemap and robots.txt together, not about any one page. It is raised once per scan with the conflicting entries listed on it.


Database Table / Prisma Model

AuditIssue (url_id null)

The blocked entries are not stored as page rows: robots.txt forbids them, so they are kept off the scan's page list. The finding is the record of them.


Finding Details

Key Type Description
blockedCount Number Distinct sitemap entries robots.txt blocks (true total)
urls Array Up to 100 { url, rule, userAgentGroup, sitemapFileUrl }
truncated Boolean True when urls is a sample
userAgent String googlebot: the section applied (falls back to *)

url is the <loc> exactly as the sitemap declares it.


Detection Dependencies

  • XML sitemap: every <loc>, with the sitemap file that listed it
  • robots.txt: fetched once per scan, per User-agent section

Further reading