Skip to content
Issue docs

Internal links to URLs blocked by robots.txt

Standardinternal_link_to_robots_blocked_urlIssue 194

What is this issue?

This page links to other pages on your site that your robots.txt tells search engines not to crawl.

A crawler that follows the link reaches a "do not enter" sign. The link is still counted as a link, so the page passes some of its authority to a URL that search engines are not allowed to read.

For a page to pass this check:

  • Every internal link on it either points at a URL search engines may crawl, or is marked rel="nofollow".

This is often deliberate. Carts, logins, account pages and checkout flows are commonly blocked and commonly linked from every page's header. That is why the check is graded STANDARD and why a nofollow link is not reported.

Example: robots.txt has Disallow: /compare, but every product page links to "Compare this product" at /compare?id=123. Each product page is reported, with the blocked target and the rule that blocks it.

Why it matters

Link authority goes to pages that cannot use it. Internal links are how a site tells search engines which pages matter. A link to a blocked URL spends some of that signal on a page that will never be crawled.

A blocked URL can still be indexed. robots.txt stops crawling, not indexing. A blocked URL with many internal links pointing at it can appear in search results as a bare address with no description ("No information is available for this page"). That is rarely what the site wants.

It can be a mistake. A Disallow written for one section often catches more than intended, such as Disallow: /p also blocking /products/. Links into a blocked area are often the first sign of it.

Why STANDARD. Much of the time the block is intentional and the cost is small. The finding is there so you can confirm it is intentional.

How to fix it

Open the finding. It lists the blocked targets and the rule that blocks each one.

  1. If the target should be in search, fix robots.txt. Remove or narrow the rule. A common cause is a prefix that matches more than intended.

    # Before: also blocks /products/ and /pricing
    Disallow: /p
    
    # After
    Disallow: /p/
  2. If the target should stay out of search, mark the links rel="nofollow". This is the usual answer for carts, logins and account pages, and it clears the finding.

    <a href="/cart" rel="nofollow">Cart</a>
  3. If the link should not be there at all, remove it.

  4. If the links are deliberate and you would rather not change them, ignore the check for this site: on the issue, choose Ignore on → Every page. The finding is raised on the pages that do the linking (often every page, through a shared header), so a URL pattern naming the target will not match it; to keep it on some sections, choose Only these pages and list the linking pages' paths instead. Ignored findings stay listed but no longer cost score, from the next scan on.

Examples

Example 1: a blocked comparison tool

robots.txt:

User-agent: *
Disallow: /compare

Problematic:

<a href="/compare?id=123">Compare this product</a>

Corrected, if the comparison pages should stay out of search:

<a href="/compare?id=123" rel="nofollow">Compare this product</a>

Example 2: a rule that blocks more than intended

robots.txt:

User-agent: *
Disallow: /p

Problematic: every link to /products/… and /pricing is reported, with the rule /p.

Corrected:

User-agent: *
Disallow: /p/

Example 3: not reported

<a href="/cart" rel="nofollow">Cart</a>
<a href="/login" rel="nofollow">Sign in</a>
<a href="https://other-site.com/admin">Partner admin</a>

The first two are nofollow. The third is an external link, which your robots.txt does not govern.

How PixyScan detects this

  1. robots.txt is fetched once per scan, and the section a search engine obeys is selected: Googlebot if the file has one, otherwise *. This is independent of the scan's "Respect robots.txt" setting, which controls what PixyScan fetches, not what Google may fetch.

  2. Every internal link on the page is tested against that section: the path and query string, * wildcards, $ end anchors, longest rule wins, Allow wins a tie. The #fragment is ignored, so /cart and /cart#top are one target.

  3. The page is reported when at least one target is blocked. The finding gives the number of distinct blocked targets and up to five of them, each with the Disallow rule that blocks it.

What is deliberately not reported:

  • Links marked rel="nofollow". The link already tells crawlers not to follow it, which is the standard way to link a page you keep out of search.
  • A site whose own home page is blocked. If robots.txt disallows the whole site, every link on every page points at a blocked URL. That is one fault in robots.txt, not one per page, so no page is reported for it.
  • External links. Your robots.txt has no say over other domains.

Note: when the scan respects robots.txt, links to blocked URLs are left out of the stored link graph, so the blocked pages never appear on the Pages screen. This finding is where those links are still visible.

What we store

Storage Level

Page Level: raised on the page that contains the links.


Database Table / Prisma Model

AuditIssue (url_id = the linking page)

Nothing else is stored for these links. When the scan respects robots.txt, links to blocked URLs are not written to the internal link graph, because storing a link would create a page row for a URL the crawl may not fetch. The finding is the only record of them.


Finding Details

Key Type Description
blockedTargetCount Number Distinct blocked targets linked from this page (true total)
blockedTargets Array Up to 5 { url, rule }: target and the blocking Disallow
truncated Boolean True when there are more than 5 targets
userAgent String googlebot: the section applied (falls back to *)

Detection Dependencies

  • robots.txt: fetched once at the start of the crawl
  • Page markup: the page's <a href> links and their rel values

Further reading