Skip to content
Issue docs

Page tells search engines not to follow its links

Importantpage_nofollow_directiveIssue 154

What is this issue?

The page carries a page-wide nofollow: an instruction to search engines not to follow any link on it.

It can be given in three places, and all three count:

  • a robots meta tag — <meta name="robots" content="nofollow">
  • a Googlebot-only meta tag — <meta name="googlebot" content="nofollow">
  • the HTTP response header — X-Robots-Tag: nofollow

It is reported only on a page that is meant to be indexed. A page that is also noindexed — noindex, nofollow, or none, which means the same — is already reported as "Pages that search engines cannot index", and is not reported again here.

This is not the same thing as rel="nofollow" on an individual link, which is the normal way to say "don't vouch for this one destination" and is never reported here.

Example: a category template ships with <meta name="robots" content="index, nofollow"> left over from a staging configuration. The category page itself is indexed, but none of the product links on it are followed — so products reachable only through the category are never discovered through it.

Why it matters

  • Discovery stops at this page. Search engines find pages by following links. A page-wide nofollow tells them not to use any link on it, so every page that is linked only from here has no route into the index through it. On a listing or category template that can be most of a catalogue.

  • Internal signals stop here too. Links are how a site tells search engines which of its pages matter. A nofollowed page passes nothing along any of its links.

  • It is usually an accident. The directive is a common leftover from a staging environment, a CMS "discourage search engines" setting, or a CDN rule written for a different path. Unlike noindex, nothing visible changes when it ships, so it can sit unnoticed for months.

  • The fix is almost always link-level. If a page genuinely has links that should not be followed — a login link, a sponsored link — rel="nofollow", rel="sponsored" or rel="ugc" on those links says so without cutting off the rest.

How to fix it

  1. Find where the directive comes from. The finding names the route: the robots meta tag in the template, or the X-Robots-Tag header set by the server or CDN.

  2. If the page's links should be followed — almost always the case — remove nofollow from the directive list. Leave the rest in place.

    <!-- Before -->
    <meta name="robots" content="index, nofollow" />
    
    <!-- After -->
    <meta name="robots" content="index, follow" />

    For the header, edit the server or CDN rule:

    # Before
    X-Robots-Tag: nofollow
    
    # After: remove the header, or
    X-Robots-Tag: max-image-preview:large
  3. If only some links should not be followed, mark those links instead:

    <a href="/login" rel="nofollow">Log in</a>
    <a href="https://partner.example/offer" rel="sponsored">Partner offer</a>
  4. If the page is meant to be hidden from search (a staging copy, a thank-you page), use noindex instead. A noindexed page is not reported here — it is reported as "Pages that search engines cannot index", where you can dismiss it if hiding it is intended.

Examples

Example 1: A followable page

<head>
  <meta name="robots" content="index, follow, max-image-preview:large" />
</head>

Passes. Nothing tells search engines not to follow the page's links.

<body>
  <a href="/login" rel="nofollow">Log in</a>
  <a href="/products">Products</a>
</body>

Passes. rel="nofollow" qualifies one link; the page's other links are followed.

Example 3: A leftover staging directive

<head>
  <meta name="robots" content="index, nofollow" />
</head>

Fails. The page is indexed, but none of its links are followed.

Example 4: Set by the CDN

HTTP/2 200
content-type: text/html
x-robots-tag: nofollow

Fails. The HTML is clean; the header applies the same directive.

Example 5: A deliberately hidden page

<!-- https://example.com/thank-you -->
<meta name="robots" content="noindex, nofollow" />

Not reported here. The page is noindexed, which is reported once as "Pages that search engines cannot index". content="none" is treated the same way.

Example 6: A value, not a directive

<meta name="robots" content="max-image-preview:none" />

Passes. none here is the value of max-image-preview, not the none directive.

How PixyScan detects this

  1. Every robots meta tag is read. All <meta name="robots"> and <meta name="googlebot"> tags, with the name matched case-insensitively, are joined into one directive list — which is how search engines combine several tags.

  2. The X-Robots-Tag header is read. The header as the server sent it, including a header sent more than once. A user-agent-scoped value such as googlebot: nofollow counts.

  3. Directive names, not substrings. Each list is split into directives and each directive's name is read, using the same parser that decides whether a page is noindexed (#21). So max-image-preview:none — a value, not the none directive — is not mistaken for one.

  4. Reported when either route declares nofollow or none — and the page is not noindexed. One finding per page, saying which route carried it (meta, header or both).

  5. Not reported by the robots meta tag check (#26) as well. That check used to list nofollow among its "harmful" directives; it no longer does, so one directive produces one finding.

  6. Link-level rel="nofollow" is ignored. It is the correct way to qualify one link and is never a page-wide directive.

  7. Noindexed pages are not reported. "Noindexed" is the same verdict as "Pages that search engines cannot index" (#21): noindex or none in the first robots meta tag or the X-Robots-Tag header. That check already reports the page, and Google has said a page left noindexed long enough is eventually treated as nofollow too — so noindex, nofollow or none on a deliberate login or thank-you page is one decision, reported once.

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageSeoBasicsData (page_seo_basics_data). The finding itself is stored on audit_issues.details.


Fields Used

Field Type Description
x_robots_tag String? The X-Robots-Tag response header verbatim; '' when the response was read and carried none, NULL when no response headers were read
indexability_noindex_absent Boolean? #21's verdict: false when the first robots meta tag or the header carries noindex or none
canonical_url String? The first <link rel="canonical"> href, as written

The robots meta content #26 reads is also stored, on PageHtmlHeadAudit (robots_meta_content) — the first name="robots" tag only.


Stored Fields on the finding

Field Type Description
robotsMeta String? Every robots and googlebot meta tag's content, joined
xRobotsTag String? The X-Robots-Tag header, or null when none was sent
source String meta, header or both — where the nofollow came from

Detection Dependencies

  • HTML Document — the robots and googlebot meta tags
  • HTTP Response Headers — the X-Robots-Tag header

Further reading