CSS or JavaScript files blocked by robots.txt
What is this issue?
Search engines do not index the raw HTML of a page. Google downloads the page's stylesheets and scripts, renders it the way a browser would, and indexes what it sees.
Every one of those downloads obeys robots.txt. If a Disallow: rule covers the folder
your CSS or JavaScript is served from, Google renders the page without it.
This check reports the stylesheets and scripts your crawled pages actually load that robots.txt blocks for search engines.
For a site to pass this check:
- Every stylesheet and script on your own host that a page loads is allowed by
robots.txt for Googlebot (or for
*, when the file has no Googlebot section).
Example: a site keeps crawlers out of its build output with Disallow: /assets/,
meaning to hide source maps. The same folder holds app.css and app.js, which every
page loads. Google indexes every page as unstyled text with no interactive content.
Why it matters
Google sees a different page from your visitors. With the stylesheet blocked, the layout collapses. With the script blocked, anything the script draws (navigation, product grids, reviews, often the main content of a single-page app) does not exist for Google.
Mobile-friendliness is judged on the rendered page. An unstyled page has text that is too small and tap targets that are too close together. Google can then treat it as not mobile-friendly, and it ranks on mobile accordingly.
Content can disappear from the index. Text that only appears once a script runs is not indexed when that script cannot be fetched.
Why IMPORTANT and not CRITICAL. The pages are still crawled and indexed. Google gets a worse version of them, but they are not missing. The fix is usually one line in one file.
Why it is one finding. The fault is a rule in robots.txt, not something wrong on any one page. PixyScan reports it once for the site and counts the pages affected, instead of charging every page for the same line.
How to fix it
Open the finding. Each blocked file is listed with the exact
Disallow:rule that blocks it.Remove or narrow that rule. If the folder holds nothing secret (and CSS and JavaScript are public by definition: anyone can download them), delete the line.
# Before User-agent: * Disallow: /assets/ # After: nothing in /assets/ needs hiding User-agent: *Or carve the files back out with
Allow. When the folder really holds something you want kept out of search, allow the stylesheet and script paths explicitly. The longer rule wins.User-agent: * Disallow: /assets/ Allow: /assets/*.css Allow: /assets/*.jsCheck the
Googlebotsection as well. If your file has aUser-agent: Googlebotsection, Google ignores the*section entirely. A rule added only to*will not help.Confirm in Search Console. Use the URL Inspection tool's "Test live URL" and look at the rendered screenshot and the list of page resources that could not be loaded.
Examples
Example 1: the build folder is disallowed
Problematic:
User-agent: *
Disallow: /static/<link rel="stylesheet" href="/static/css/main.4f2a.css">
<script src="/static/js/main.9c1e.js" defer></script>Both files are blocked. Google renders every page without styles or scripts.
Corrected:
User-agent: *
Disallow: /static/maps/Only the source maps stay hidden; the CSS and JavaScript load.
Example 2: a query-string rule catches versioned assets
Problematic: a rule meant for faceted search also matches cache-busting parameters.
User-agent: *
Disallow: /*?ver=<script src="/wp-includes/js/jquery/jquery.min.js?ver=3.7.1"></script>Corrected: scope the rule to the URLs it was written for.
User-agent: *
Disallow: /shop/*?ver=Example 3: what passes
User-agent: *
Disallow: /admin/
User-agent: SemrushBot
Disallow: /<link rel="stylesheet" href="https://example.com/assets/site.css">
<script src="https://cdn.other-vendor.com/widget.js"></script>/assets/ is allowed for search engines. The Disallow: / applies to SemrushBot only.
The CDN script answers to the CDN's own robots.txt, so this check does not judge it.
How PixyScan detects this
This check runs once, after the crawl, using two things the crawl recorded.
The files each page loads. While crawling, PixyScan reads every page's markup and records the stylesheets (
<link rel="stylesheet">,as="style"preloads) and scripts (<script src>,modulepreload) it references, up to the first 100 per page. This happens on every plan. Only weighing those files uses your plan's request budget.Your robots.txt, section by section. The file is fetched once per scan and stored with each
User-agentsection kept separate.The section a search engine obeys is selected: the
Googlebotsection if your file has one, otherwise the*section. ADisallow: /aimed at some other bot does not count.Each distinct stylesheet and script on your own host is tested against that section, using the standard robots.txt rules:
- the path and query string are matched (
Disallow: /*?ver=blocksjquery.js?ver=3.7); *matches anything and a trailing$anchors the end;- the longest matching rule wins, and
Allowwins a tie.
- the path and query string are matched (
One finding is raised for the site if any file is blocked. It lists up to 100 of them, each with the rule that blocks it, how many pages load it, and one example page. It also gives the total number of blocked files and the number of pages affected.
What is never reported:
- Files on another host, such as a CDN or a third-party tag. That host has its own
robots.txt, which this scan did not read. Your site's host and its
www.twin both count as your own host. - Images and fonts. Images are fetched by a different crawler, and neither changes how Google lays out and runs the page the way CSS and JavaScript do.
- Scans that stored no per-section robots.txt rules (older scans). "We did not look" is not reported as either a pass or a failure.
What we store
Storage Level
Site Level: the finding is about robots.txt, not about any one page, so it is raised once per scan with the blocked files listed on it.
Database Table / Prisma Model
AuditIssue (url_id null), computed from PageResource and
SiteCrawlBehaviourData.
Fields Used
| Field | Type | Description |
|---|---|---|
| PageResource.resourceUrl | String | Absolute address of a stylesheet or script a page references |
| PageResource.resourceType | String | css or js (font rows are ignored by this check) |
| PageResource.urlId | String | The page that references it, for the page counts |
| SiteCrawlBehaviourData.robotsTxtData | Json | groups[]: each User-agent section with its allow / disallow |
| Scan.baseUrl / Site.url | String? | The host the robots.txt belongs to |
Finding Details
| Key | Description |
|---|---|
blockedCount |
Distinct blocked files (true total) |
cssCount, jsCount |
The same total, split by type |
affectedPageCount |
Distinct crawled pages that load at least one blocked file |
userAgentGroup |
googlebot or *: the section that was applied |
rules |
The distinct Disallow patterns responsible |
resources |
Up to 100 { url, type, rule, pageCount, examplePage } |
truncated |
True when resources is a sample |
Detection Dependencies
- Page markup: stylesheet and script references, read on every crawled page
- robots.txt: fetched once per scan, stored per
User-agentsection - The whole crawl: the check runs after the last page is stored