robots.txt is served but is not valid
What is this issue?
This issue fires when {origin}/robots.txt is served, with content in it, and that content still does not configure crawling. Two shapes of that are recorded under status:
no-directives— the file has content, but nothing in it parses as a directive. Comments only, or prose, or an HTML error page returned with HTTP 200.syntax-errors— the file does parse, and some of its lines do not. Those lines are listed insyntaxErrors, each naming its line number.
A line counts as unusable when a crawler reading the file would discard it:
- there is no
:separating a directive from its value - the directive name is not one crawlers recognise —
Disalow: /adminis the classic User-agent:is present with no value, so it opens no groupAllow,DisalloworCrawl-delayappears before anyUser-agentline, or after a blank line has closed the group it belonged to, so it applies to no crawlerCrawl-delayis not a numberSitemapis not an absolutehttp(s)://URL — the specification requires one
This is the counterpart to No robots.txt file on the site, which covers the case where there is no file to read at all — a 404, a 5xx, or a zero-length body. The two are separate because the fix is: correct the file you have, rather than publish one you do not.
A robots.txt that is valid but restrictive — Disallow: / for every crawler, say — is neither of these issues. It is a file that says exactly what it means, and PixyScan reports the parsed rules in the crawl behaviour analysis rather than as a defect.
Why it matters
The damage here is that the file looks fine. Someone wrote rules, deployed them, and reasonably believes they are in force — but a crawler reading the file throws those lines away without complaint. There is no error page, no warning in Search Console, and nothing in the file itself to suggest anything is wrong.
What that costs:
- Exclusions that never applied.
Disalow: /adminis one transposed letter, and it means the admin path has been open to crawlers for as long as the typo has been live. The same goes for aDisallowthat sits above the firstUser-agentline, or below a blank line that closed the group it was meant to join. - A sitemap crawlers cannot follow.
Sitemap: /sitemap.xmlis discarded; the directive requires an absolute URL. Discovery falls back to whatever links the crawler happens to find. - A crawl rate you did not get. A non-numeric
Crawl-delayis ignored, so a crawler you meant to slow down carries on at its own pace. - Nothing configured at all, in the
no-directivescase. The effect is the same as having no robots.txt, with the added problem that the file's presence hides it.
It is important rather than critical because the failure mode is permissive: a crawler that discards a line crawls more than you intended, rather than less. Nothing gets deindexed by this. But every rule you wrote is worth exactly as much as the crawler's ability to read it, and here the answer is nothing.
How to fix it
Read
syntaxErrorsin the finding. Every unusable line is listed with its line number and what is wrong with it. Open/robots.txtand go to those lines.Fix the line, not the file. The common causes and their corrections:
What the finding says What to change "..." is not a directive — no ":" separatorThe line is prose or a stray path. Delete it, or comment it out with #."disalow" is not a robots.txt directiveA misspelling. Only User-agent,Allow,Disallow,Crawl-delayandSitemapare widely honoured.disallow appears outside a User-agent groupMove the rule below a User-agent:line — and check for a blank line above it, since a blank line ends the group.User-agent has no valueGive it one: User-agent: *.Crawl-delay value "..." is not a numberUse seconds as a number: Crawl-delay: 10.Sitemap "..." is not an absolute http(s) URLWrite it in full: Sitemap: https://example.com/sitemap.xml.If
statusisno-directives, check what is actually being served. A file of comments needs at least one real directive. A file of HTML usually means the server is returning an error page or an SPA shell with HTTP 200 for a path it does not have; fix the routing so/robots.txtserves a plain-text file.Mind the blank lines. A blank line ends the current
User-agentgroup. This is valid:User-agent: * Disallow: /admin/ Sitemap: https://example.com/sitemap.xmland this quietly is not, because the blank line orphans the
Disallow:User-agent: * Disallow: /admin/Verify. Fetch
/robots.txtyourself and confirm it comes back astext/plain, then check the rules you care about in Google's robots.txt report before re-running the scan.
Examples
Example 1: A misspelled directive
Scenario: The admin area was meant to be closed to crawlers a year ago.
Problematic State (Fails):
User-agent: *
Disalow: /admin/
Sitemap: https://example.com/sitemap.xmldisalow is not a directive a crawler recognises, so the line is discarded and /admin/ has been open the whole time. The file still parses — there is a User-agent group and a Sitemap — so the finding comes back as status: syntax-errors with one entry:
Line 2: "disalow" is not a robots.txt directive, so the line is ignored.Corrected State (Passes):
User-agent: *
Disallow: /admin/
Sitemap: https://example.com/sitemap.xmlExample 2: A blank line that closed the group
Scenario: Someone added spacing to make the file easier to read.
Problematic State (Fails):
User-agent: *
Disallow: /admin/
Disallow: /checkout/
Sitemap: /sitemap.xmlTwo separate faults. The blank line on line 2 ends the User-agent group, so both Disallow rules sit outside any group and apply to no crawler. And Sitemap needs an absolute URL — a path is discarded. Reported as status: syntax-errors:
Line 3: disallow appears outside a User-agent group, so it applies to no crawler.
Line 4: disallow appears outside a User-agent group, so it applies to no crawler.
Line 6: Sitemap "/sitemap.xml" is not an absolute http(s) URL.Corrected State (Passes):
User-agent: *
Disallow: /admin/
Disallow: /checkout/
Sitemap: https://example.com/sitemap.xmlThe blank line before Sitemap is fine — Sitemap is not a group directive, so it does not need to sit inside one.
Example 3: A file that configures nothing
Scenario: /robots.txt returns HTTP 200, and the response body is the site's SPA shell because the router has no handler for that path.
Problematic State (Fails):
<!DOCTYPE html>
<html>
<head><title>Example</title></head>
<body><div id="root"></div></body>
</html>Nothing here parses as a directive: no User-agent rule, no Sitemap. This is the status: no-directives branch, and the same applies to a file containing only comments, or only prose. The presence of the file is what hides the problem — the site looks configured and is not.
Corrected State (Passes): Fix the routing so /robots.txt is served as a plain-text file, with at least one real directive in it:
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xmlHow PixyScan detects this
Fetch. The crawler requests
{origin}/robots.txtonce per scan. This check only runs when the file came back with HTTP 200 — anything else is No robots.txt file on the site, not this issue.Content gate. A zero-length or whitespace-only body stops here and raises the missing issue instead. There is no configuration to correct, so calling it misconfigured would point you at content that does not exist.
Line walk. The remaining body is walked line by line, the same way the parser walks it:
#comments stripped, blank lines ending the currentUser-agentgroup. A line is recorded as an error if and only if the parser would discard it — no:separator, an unrecognised directive name, a valuelessUser-agent, a rule outside any group, a non-numericCrawl-delay, or aSitemapthat is not an absolutehttp(s)URL. Each error names its line number.Classification. If the file yields no
User-agentrule and noSitemapdirective at all, the finding is raised withstatus: 'no-directives'. Otherwise, if any line failed the walk, it is raised withstatus: 'syntax-errors'. A file that parses cleanly raises nothing.One row per scan. Both branches carry the full list in
details.syntaxErrorsand the count indetails.syntaxErrorCount, and the message quotes the first error so the issue list is actionable without opening the detail page. A file that satisfies both branches produces a single finding.Scope. This is a site-level check — one finding per scan, not one per URL.
What we store
Storage Level
Site Level — one row per scan, keyed by scanId, not per URL. The finding is written with urlId NULL.
Database Table / Prisma Model
SiteCrawlBehaviourData — the parsed robots.txt payload.
AuditIssue — the finding itself, with the unusable lines in details.
Stored Fields
SiteCrawlBehaviourData
| Field | Type | Description |
|---|---|---|
| robotsTxtData | Json? | The robots.txt payload. raw is the file byte for byte — original casing, comments, blank lines and directive order — which is what makes the line numbers in the finding resolvable. rawTruncated marks a file cut at 512k chars. userAgents, disallowedPaths, allowedPaths, crawlDelays and groups are the parsed rules; a line this issue reports is, by definition, absent from them. null when the file was not fetched or came back empty. |
AuditIssue
| Field | Type | Description |
|---|---|---|
| issueCode | String | robots_txt_misconfigured |
| scanId | String | The scan this finding belongs to |
| urlId | String? | NULL — the check is site-level |
| details | Json? | status (no-directives or syntax-errors), syntaxErrors (every unusable line, each prefixed with its line number), syntaxErrorCount, and message (quotes the first error so the issue list is actionable without opening the detail page) |
The per-line list lives only in details. The crawler also computes a syntaxValidation block alongside robotsTxtData, but no DTO declares it, so the global validation pipe strips it before it reaches Prisma — details.syntaxErrors is the durable copy.
Detection Dependencies
- The following data sources are required to evaluate this issue:
- robots.txt — the raw text of
{origin}/robots.txt, needed line by line rather than parsed, since the lines this issue reports are exactly the ones parsing discards - HTTP Response — the status code of that fetch; the check only runs on HTTP 200