Page tells search engines not to follow its links
What is this issue?
The page carries a page-wide nofollow: an instruction to search engines not to
follow any link on it.
It can be given in three places, and all three count:
- a robots meta tag —
<meta name="robots" content="nofollow"> - a Googlebot-only meta tag —
<meta name="googlebot" content="nofollow"> - the HTTP response header —
X-Robots-Tag: nofollow
It is reported only on a page that is meant to be indexed. A page that is also
noindexed — noindex, nofollow, or none, which means the same — is already reported
as "Pages that search engines cannot index", and is not reported again here.
This is not the same thing as rel="nofollow" on an individual link, which is the
normal way to say "don't vouch for this one destination" and is never reported here.
Example: a category template ships with <meta name="robots" content="index, nofollow">
left over from a staging configuration. The category page itself is indexed, but none
of the product links on it are followed — so products reachable only through the
category are never discovered through it.
Why it matters
Discovery stops at this page. Search engines find pages by following links. A page-wide nofollow tells them not to use any link on it, so every page that is linked only from here has no route into the index through it. On a listing or category template that can be most of a catalogue.
Internal signals stop here too. Links are how a site tells search engines which of its pages matter. A nofollowed page passes nothing along any of its links.
It is usually an accident. The directive is a common leftover from a staging environment, a CMS "discourage search engines" setting, or a CDN rule written for a different path. Unlike
noindex, nothing visible changes when it ships, so it can sit unnoticed for months.The fix is almost always link-level. If a page genuinely has links that should not be followed — a login link, a sponsored link —
rel="nofollow",rel="sponsored"orrel="ugc"on those links says so without cutting off the rest.
How to fix it
Find where the directive comes from. The finding names the route: the robots meta tag in the template, or the
X-Robots-Tagheader set by the server or CDN.If the page's links should be followed — almost always the case — remove
nofollowfrom the directive list. Leave the rest in place.<!-- Before --> <meta name="robots" content="index, nofollow" /> <!-- After --> <meta name="robots" content="index, follow" />For the header, edit the server or CDN rule:
# Before X-Robots-Tag: nofollow # After: remove the header, or X-Robots-Tag: max-image-preview:largeIf only some links should not be followed, mark those links instead:
<a href="/login" rel="nofollow">Log in</a> <a href="https://partner.example/offer" rel="sponsored">Partner offer</a>If the page is meant to be hidden from search (a staging copy, a thank-you page), use
noindexinstead. A noindexed page is not reported here — it is reported as "Pages that search engines cannot index", where you can dismiss it if hiding it is intended.
Examples
Example 1: A followable page
<head>
<meta name="robots" content="index, follow, max-image-preview:large" />
</head>Passes. Nothing tells search engines not to follow the page's links.
Example 2: A link-level nofollow
<body>
<a href="/login" rel="nofollow">Log in</a>
<a href="/products">Products</a>
</body>Passes. rel="nofollow" qualifies one link; the page's other links are followed.
Example 3: A leftover staging directive
<head>
<meta name="robots" content="index, nofollow" />
</head>Fails. The page is indexed, but none of its links are followed.
Example 4: Set by the CDN
HTTP/2 200
content-type: text/html
x-robots-tag: nofollowFails. The HTML is clean; the header applies the same directive.
Example 5: A deliberately hidden page
<!-- https://example.com/thank-you -->
<meta name="robots" content="noindex, nofollow" />Not reported here. The page is noindexed, which is reported once as "Pages that
search engines cannot index". content="none" is treated the same way.
Example 6: A value, not a directive
<meta name="robots" content="max-image-preview:none" />Passes. none here is the value of max-image-preview, not the none directive.
How PixyScan detects this
Every robots meta tag is read. All
<meta name="robots">and<meta name="googlebot">tags, with thenamematched case-insensitively, are joined into one directive list — which is how search engines combine several tags.The X-Robots-Tag header is read. The header as the server sent it, including a header sent more than once. A user-agent-scoped value such as
googlebot: nofollowcounts.Directive names, not substrings. Each list is split into directives and each directive's name is read, using the same parser that decides whether a page is noindexed (#21). So
max-image-preview:none— a value, not thenonedirective — is not mistaken for one.Reported when either route declares
nofollowornone— and the page is not noindexed. One finding per page, saying which route carried it (meta,headerorboth).Not reported by the robots meta tag check (#26) as well. That check used to list
nofollowamong its "harmful" directives; it no longer does, so one directive produces one finding.Link-level
rel="nofollow"is ignored. It is the correct way to qualify one link and is never a page-wide directive.Noindexed pages are not reported. "Noindexed" is the same verdict as "Pages that search engines cannot index" (#21):
noindexornonein the first robots meta tag or theX-Robots-Tagheader. That check already reports the page, and Google has said a page left noindexed long enough is eventually treated as nofollow too — sonoindex, nofollowornoneon a deliberate login or thank-you page is one decision, reported once.
What we store
Storage Level
Page Level
Database Table / Prisma Model
PageSeoBasicsData (page_seo_basics_data). The finding itself is stored on
audit_issues.details.
Fields Used
| Field | Type | Description |
|---|---|---|
| x_robots_tag | String? | The X-Robots-Tag response header verbatim; '' when the response was read and carried none, NULL when no response headers were read |
| indexability_noindex_absent | Boolean? | #21's verdict: false when the first robots meta tag or the header carries noindex or none |
| canonical_url | String? | The first <link rel="canonical"> href, as written |
The robots meta content #26 reads is also stored, on PageHtmlHeadAudit
(robots_meta_content) — the first name="robots" tag only.
Stored Fields on the finding
| Field | Type | Description |
|---|---|---|
| robotsMeta | String? | Every robots and googlebot meta tag's content, joined |
| xRobotsTag | String? | The X-Robots-Tag header, or null when none was sent |
| source | String | meta, header or both — where the nofollow came from |
Detection Dependencies
- HTML Document — the robots and googlebot meta tags
- HTTP Response Headers — the X-Robots-Tag header