Skip to content
Issue docs

Robots meta tag and X-Robots-Tag header disagree

Standardrobots_directives_conflictIssue 156

What is this issue?

The page states its robots rules in two places — a robots meta tag in the HTML and an X-Robots-Tag HTTP header — and the two take opposite sides:

  • one says noindex and the other says index (or all), or
  • one says nofollow and the other says follow (or all).

none counts as both noindex and nofollow.

A route that does not mention indexing or following has not taken a side. A header saying noindex beside a meta tag that only sets preview directives is not a disagreement — the noindex itself is reported as "Pages that search engines cannot index".

Example: the template says <meta name="robots" content="index, follow">, and a CDN rule written for the staging hostname also matches production and adds X-Robots-Tag: noindex.

Why it matters

  • The restrictive rule wins, silently. Search engines resolve a conflict by obeying the more restrictive directive. A page whose template says index and whose server says noindex is not indexed, and nothing in the HTML shows why.

  • Two places, two owners. The meta tag lives in the template; the header lives in the server or CDN configuration. Whoever edits one usually looks only at that one, so the conflict survives every attempt to fix it from the wrong side.

  • It is a reliable sign of a misapplied rule. Conflicts usually come from a header rule scoped too widely, or a template default nobody updated.

How to fix it

  1. Decide the rule the page should have. Usually the template is right and a header rule is applying where it should not, but check.

  2. State it in one place. The simplest fix is to remove the directive from the route that should not be setting it.

    # Before
    <meta name="robots" content="index, follow">
    X-Robots-Tag: noindex
    
    # After (page should be indexed): remove the header rule for this path
    <meta name="robots" content="index, follow">
  3. If both routes must stay, make them say the same thing.

  4. Check the scope of header rules. A header added at the CDN or web server for a staging host, a file type or a path prefix often matches more than intended.

Examples

Example 1: The two routes agree

<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow

Passes. Redundant, but consistent.

Example 2: One route is silent

<meta name="robots" content="max-snippet:-1, max-image-preview:large">
X-Robots-Tag: noindex

Passes this check. The meta tag says nothing about indexing. The noindex is reported separately as "Pages that search engines cannot index".

Example 3: Opposite indexing rules

<meta name="robots" content="index, follow">
X-Robots-Tag: noindex

Fails. The template asks for indexing, the server forbids it; the page is not indexed.

Example 4: Opposite on both dimensions

<meta name="robots" content="all">
X-Robots-Tag: none

Fails on both indexing and following.

How PixyScan detects this

  1. Both routes must be present. Every <meta name="robots"> and <meta name="googlebot"> tag (name matched case-insensitively) joined into one list, and a non-empty X-Robots-Tag header. With only one route there is nothing to disagree with.

  2. Each route's stance is read on two dimensions.

    • Indexing: noindex if the route carries noindex or none; otherwise index if it explicitly says index or all; otherwise it has not said.
    • Following: nofollow if it carries nofollow or none; otherwise follow if it explicitly says follow or all; otherwise it has not said.

    The restrictive half is read with the same parser #21 uses, so a value such as max-image-preview:none is never mistaken for none. A user-agent scope on the header (googlebot: noindex) is read as applying.

  3. Reported when both routes have a stance on a dimension and they differ. One finding per page, listing each conflicting dimension with what each route said.

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageSeoBasicsData (page_seo_basics_data). The finding itself is stored on audit_issues.details.


Fields Used

Field Type Description
x_robots_tag String? The X-Robots-Tag response header verbatim; '' when the response was read and carried none, NULL when no response headers were read
indexability_noindex_absent Boolean? #21's verdict: false when the first robots meta tag or the header carries noindex or none
canonical_url String? The first <link rel="canonical"> href, as written

The robots meta content #26 reads is also stored, on PageHtmlHeadAudit (robots_meta_content) — the first name="robots" tag only.


Stored Fields on the finding

Field Type Description
robotsMeta String Every robots and googlebot meta tag's content, joined
xRobotsTag String The X-Robots-Tag header
conflicts Array One entry per conflicting dimension: { dimension, meta, header }

Detection Dependencies

  • HTML Document — the robots and googlebot meta tags
  • HTTP Response Headers — the X-Robots-Tag header

Further reading