Skip to content
Issue docs

Sitemap lists pages that carry a noindex directive

Standardsitemap_lists_noindex_urlIssue 147

What is this issue?

Listing a page in a sitemap asks search engines to index it. A noindex directive on that page asks them not to. This check reports the pages where a site is saying both things at once.

For a sitemap to pass this check:

  • No URL it lists carries a noindex, in either the robots meta tag or the X-Robots-Tag response header.

A noindex is not itself a fault — thank-you pages, filtered listings, internal search results and staging locales are all legitimately hidden. What this check reports is advertising them in the file whose job is to say what should be indexed.

Example: a site adds noindex to its tag-archive pages to stop them competing with the articles they list. The sitemap is generated from the CMS's full page list and still includes all four hundred of them.

Why it matters

The two signals contradict each other, and the sitemap loses. Google honours the noindex. The entry in the sitemap achieves nothing except a fetch.

Crawl budget goes to pages that will never appear. On a large site this is the single biggest source of wasted sitemap requests, because the pages involved — filtered views, tag archives, paginated listings — are usually numerous.

Search Console reports it back as a coverage problem. "Excluded by 'noindex' tag" on pages you submitted is a report a site owner has to interpret every time it appears.

Why this is graded STANDARD. The indexability of the page itself is already reported, and scored, by the check that owns that fault. What this one adds is that your discovery file disagrees with it — real, but indirect, with a cheap fix. Charging it at full weight would bill one page twice for one decision.

Fixing it makes the sitemap a statement about what you want indexed, which is what makes it worth submitting.

How to fix it

  1. Read the list on the finding — each entry names the address as the sitemap declares it and the file it came from.

  2. Decide which signal you meant. For each page:

    • if it is deliberately hidden, remove it from the sitemap;
    • if it is meant to rank, remove the noindex.
  3. Fix the generator, not the file. These entries appear because the sitemap is built from "every page in the CMS" rather than "every page we want indexed". Filtering the generator on the same flag that sets the noindex is what keeps them in step permanently.

  4. Check the X-Robots-Tag header too. A noindex set at the CDN or in server configuration does not appear in the page's markup, so a template-level audit will not see it — but search engines treat it identically.

  5. Re-submit the sitemap so the coverage report reflects the change.

Examples

Example 1 — tag archives in the sitemap

Problematic: the page is hidden, and advertised.

<url><loc>https://example.com/tag/summer</loc></url>
<!-- https://example.com/tag/summer -->
<meta name="robots" content="noindex, follow" />

Corrected: drop the entry from the sitemap and leave the noindex in place. The archive still passes link equity through — that is what follow is for — it simply is not something you are asking to have indexed.


Example 2 — a noindex set at the CDN

Problematic: nothing in the markup looks wrong, because the directive is a header.

GET /reports/internal-q3.pdf
HTTP/1.1 200 OK
X-Robots-Tag: noindex
<url><loc>https://example.com/reports/internal-q3.pdf</loc></url>

Corrected: remove the sitemap entry. A file served with X-Robots-Tag: noindex is one you have decided should not appear in results.


Example 3 — the page was meant to rank

Problematic: a noindex left behind from a staging deploy.

<!-- https://example.com/pricing -->
<meta name="robots" content="noindex" />

Corrected: remove the directive. Here the sitemap was right and the page was wrong, which is the other way this finding resolves.

How PixyScan detects this

The sitemap knows what it lists; only a fetch of each page reveals its directives. So this runs once, after the crawl.

  1. Take every URL the sitemap listed, whether or not the crawl also reached it by following links.

  2. Keep only the ones that answered HTTP 200. A page that did not serve is a different finding, and it has no directives to read anyway.

  3. Read the indexability the crawl recorded — whether a noindex was present in the robots meta tag or in the X-Robots-Tag response header. Both count; Google treats them identically.

  4. Report every listed address that carries one, naming it exactly as the sitemap declared it, together with the sitemap file it came from.

What is deliberately never reported:

  • A page whose directives were never measured. A listed address the crawl could not reach, or one with no page analysis recorded — a PDF, a feed, a page that failed before analysis — is left alone rather than assumed clean or assumed noindexed.
  • A URL the sitemap does not list. A noindexed page that is not in the sitemap is not this finding at all; it is simply a hidden page, which is a normal thing for a site to have.

What we store

Storage Level

Site Level — the finding is about the sitemap, not about any one page in it, so it is raised once per scan with the offending entries listed on it.


Database Table / Prisma Model

AuditIssue (url_id null), read from Url joined against PageSeoBasicsData.


Fields Used

Field Type Description
Url.source Enum SITEMAP or BOTH — the sitemap listed this address
Url.sitemapLoc String? The <loc> verbatim, which is what the finding names
Url.statusCode Int? Only pages that answered 200 are judged
PageSeoBasicsData.indexabilityNoindexAbsent Boolean? False when a noindex directive is present
PageSeoBasicsData.xRobotsTag String? The header form of the directive, as sent

Detection Dependencies

  • XML Sitemap — the <loc> list
  • HTML Document — the meta name="robots" tag
  • HTTP Response — the X-Robots-Tag header

Further reading