Internal links to URLs blocked by robots.txt
What is this issue?
This page links to other pages on your site that your robots.txt tells search engines not to crawl.
A crawler that follows the link reaches a "do not enter" sign. The link is still counted as a link, so the page passes some of its authority to a URL that search engines are not allowed to read.
For a page to pass this check:
- Every internal link on it either points at a URL search engines may crawl, or is marked
rel="nofollow".
This is often deliberate. Carts, logins, account pages and checkout flows are commonly
blocked and commonly linked from every page's header. That is why the check is graded
STANDARD and why a nofollow link is not reported.
Example: robots.txt has Disallow: /compare, but every product page links to
"Compare this product" at /compare?id=123. Each product page is reported, with the
blocked target and the rule that blocks it.
Why it matters
Link authority goes to pages that cannot use it. Internal links are how a site tells search engines which pages matter. A link to a blocked URL spends some of that signal on a page that will never be crawled.
A blocked URL can still be indexed. robots.txt stops crawling, not indexing. A blocked URL with many internal links pointing at it can appear in search results as a bare address with no description ("No information is available for this page"). That is rarely what the site wants.
It can be a mistake. A Disallow written for one section often catches more than
intended, such as Disallow: /p also blocking /products/. Links into a blocked area
are often the first sign of it.
Why STANDARD. Much of the time the block is intentional and the cost is small. The finding is there so you can confirm it is intentional.
How to fix it
Open the finding. It lists the blocked targets and the rule that blocks each one.
If the target should be in search, fix robots.txt. Remove or narrow the rule. A common cause is a prefix that matches more than intended.
# Before: also blocks /products/ and /pricing Disallow: /p # After Disallow: /p/If the target should stay out of search, mark the links
rel="nofollow". This is the usual answer for carts, logins and account pages, and it clears the finding.<a href="/cart" rel="nofollow">Cart</a>If the link should not be there at all, remove it.
If the links are deliberate and you would rather not change them, ignore the check for this site: on the issue, choose Ignore on → Every page. The finding is raised on the pages that do the linking (often every page, through a shared header), so a URL pattern naming the target will not match it; to keep it on some sections, choose Only these pages and list the linking pages' paths instead. Ignored findings stay listed but no longer cost score, from the next scan on.
Examples
Example 1: a blocked comparison tool
robots.txt:
User-agent: *
Disallow: /compareProblematic:
<a href="/compare?id=123">Compare this product</a>Corrected, if the comparison pages should stay out of search:
<a href="/compare?id=123" rel="nofollow">Compare this product</a>Example 2: a rule that blocks more than intended
robots.txt:
User-agent: *
Disallow: /pProblematic: every link to /products/… and /pricing is reported, with the rule
/p.
Corrected:
User-agent: *
Disallow: /p/Example 3: not reported
<a href="/cart" rel="nofollow">Cart</a>
<a href="/login" rel="nofollow">Sign in</a>
<a href="https://other-site.com/admin">Partner admin</a>The first two are nofollow. The third is an external link, which your robots.txt does
not govern.
How PixyScan detects this
robots.txt is fetched once per scan, and the section a search engine obeys is selected:
Googlebotif the file has one, otherwise*. This is independent of the scan's "Respect robots.txt" setting, which controls what PixyScan fetches, not what Google may fetch.Every internal link on the page is tested against that section: the path and query string,
*wildcards,$end anchors, longest rule wins,Allowwins a tie. The#fragmentis ignored, so/cartand/cart#topare one target.The page is reported when at least one target is blocked. The finding gives the number of distinct blocked targets and up to five of them, each with the
Disallowrule that blocks it.
What is deliberately not reported:
- Links marked
rel="nofollow". The link already tells crawlers not to follow it, which is the standard way to link a page you keep out of search. - A site whose own home page is blocked. If robots.txt disallows the whole site, every link on every page points at a blocked URL. That is one fault in robots.txt, not one per page, so no page is reported for it.
- External links. Your robots.txt has no say over other domains.
Note: when the scan respects robots.txt, links to blocked URLs are left out of the stored link graph, so the blocked pages never appear on the Pages screen. This finding is where those links are still visible.
What we store
Storage Level
Page Level: raised on the page that contains the links.
Database Table / Prisma Model
AuditIssue (url_id = the linking page)
Nothing else is stored for these links. When the scan respects robots.txt, links to blocked URLs are not written to the internal link graph, because storing a link would create a page row for a URL the crawl may not fetch. The finding is the only record of them.
Finding Details
| Key | Type | Description |
|---|---|---|
blockedTargetCount |
Number | Distinct blocked targets linked from this page (true total) |
blockedTargets |
Array | Up to 5 { url, rule }: target and the blocking Disallow |
truncated |
Boolean | True when there are more than 5 targets |
userAgent |
String | googlebot: the section applied (falls back to *) |
Detection Dependencies
- robots.txt: fetched once at the start of the crawl
- Page markup: the page's
<a href>links and theirrelvalues