Skip to content
Issue docs

robots.txt is served but is not valid

Importantrobots_txt_misconfiguredIssue 6

What is this issue?

This issue fires when {origin}/robots.txt is served, with content in it, and that content still does not configure crawling. Two shapes of that are recorded under status:

  • no-directives — the file has content, but nothing in it parses as a directive. Comments only, or prose, or an HTML error page returned with HTTP 200.
  • syntax-errors — the file does parse, and some of its lines do not. Those lines are listed in syntaxErrors, each naming its line number.

A line counts as unusable when a crawler reading the file would discard it:

  • there is no : separating a directive from its value
  • the directive name is not one crawlers recognise — Disalow: /admin is the classic
  • User-agent: is present with no value, so it opens no group
  • Allow, Disallow or Crawl-delay appears before any User-agent line, or after a blank line has closed the group it belonged to, so it applies to no crawler
  • Crawl-delay is not a number
  • Sitemap is not an absolute http(s):// URL — the specification requires one

This is the counterpart to No robots.txt file on the site, which covers the case where there is no file to read at all — a 404, a 5xx, or a zero-length body. The two are separate because the fix is: correct the file you have, rather than publish one you do not.

A robots.txt that is valid but restrictive — Disallow: / for every crawler, say — is neither of these issues. It is a file that says exactly what it means, and PixyScan reports the parsed rules in the crawl behaviour analysis rather than as a defect.

Why it matters

The damage here is that the file looks fine. Someone wrote rules, deployed them, and reasonably believes they are in force — but a crawler reading the file throws those lines away without complaint. There is no error page, no warning in Search Console, and nothing in the file itself to suggest anything is wrong.

What that costs:

  • Exclusions that never applied. Disalow: /admin is one transposed letter, and it means the admin path has been open to crawlers for as long as the typo has been live. The same goes for a Disallow that sits above the first User-agent line, or below a blank line that closed the group it was meant to join.
  • A sitemap crawlers cannot follow. Sitemap: /sitemap.xml is discarded; the directive requires an absolute URL. Discovery falls back to whatever links the crawler happens to find.
  • A crawl rate you did not get. A non-numeric Crawl-delay is ignored, so a crawler you meant to slow down carries on at its own pace.
  • Nothing configured at all, in the no-directives case. The effect is the same as having no robots.txt, with the added problem that the file's presence hides it.

It is important rather than critical because the failure mode is permissive: a crawler that discards a line crawls more than you intended, rather than less. Nothing gets deindexed by this. But every rule you wrote is worth exactly as much as the crawler's ability to read it, and here the answer is nothing.

How to fix it

  1. Read syntaxErrors in the finding. Every unusable line is listed with its line number and what is wrong with it. Open /robots.txt and go to those lines.

  2. Fix the line, not the file. The common causes and their corrections:

    What the finding says What to change
    "..." is not a directive — no ":" separator The line is prose or a stray path. Delete it, or comment it out with #.
    "disalow" is not a robots.txt directive A misspelling. Only User-agent, Allow, Disallow, Crawl-delay and Sitemap are widely honoured.
    disallow appears outside a User-agent group Move the rule below a User-agent: line — and check for a blank line above it, since a blank line ends the group.
    User-agent has no value Give it one: User-agent: *.
    Crawl-delay value "..." is not a number Use seconds as a number: Crawl-delay: 10.
    Sitemap "..." is not an absolute http(s) URL Write it in full: Sitemap: https://example.com/sitemap.xml.
  3. If status is no-directives, check what is actually being served. A file of comments needs at least one real directive. A file of HTML usually means the server is returning an error page or an SPA shell with HTTP 200 for a path it does not have; fix the routing so /robots.txt serves a plain-text file.

  4. Mind the blank lines. A blank line ends the current User-agent group. This is valid:

    User-agent: *
    Disallow: /admin/
    
    Sitemap: https://example.com/sitemap.xml

    and this quietly is not, because the blank line orphans the Disallow:

    User-agent: *
    
    Disallow: /admin/
  5. Verify. Fetch /robots.txt yourself and confirm it comes back as text/plain, then check the rules you care about in Google's robots.txt report before re-running the scan.

Examples

Example 1: A misspelled directive

Scenario: The admin area was meant to be closed to crawlers a year ago.

Problematic State (Fails):

User-agent: *
Disalow: /admin/
Sitemap: https://example.com/sitemap.xml

disalow is not a directive a crawler recognises, so the line is discarded and /admin/ has been open the whole time. The file still parses — there is a User-agent group and a Sitemap — so the finding comes back as status: syntax-errors with one entry:

Line 2: "disalow" is not a robots.txt directive, so the line is ignored.

Corrected State (Passes):

User-agent: *
Disallow: /admin/
Sitemap: https://example.com/sitemap.xml

Example 2: A blank line that closed the group

Scenario: Someone added spacing to make the file easier to read.

Problematic State (Fails):

User-agent: *

Disallow: /admin/
Disallow: /checkout/

Sitemap: /sitemap.xml

Two separate faults. The blank line on line 2 ends the User-agent group, so both Disallow rules sit outside any group and apply to no crawler. And Sitemap needs an absolute URL — a path is discarded. Reported as status: syntax-errors:

Line 3: disallow appears outside a User-agent group, so it applies to no crawler.
Line 4: disallow appears outside a User-agent group, so it applies to no crawler.
Line 6: Sitemap "/sitemap.xml" is not an absolute http(s) URL.

Corrected State (Passes):

User-agent: *
Disallow: /admin/
Disallow: /checkout/

Sitemap: https://example.com/sitemap.xml

The blank line before Sitemap is fine — Sitemap is not a group directive, so it does not need to sit inside one.


Example 3: A file that configures nothing

Scenario: /robots.txt returns HTTP 200, and the response body is the site's SPA shell because the router has no handler for that path.

Problematic State (Fails):

<!DOCTYPE html>
<html>
  <head><title>Example</title></head>
  <body><div id="root"></div></body>
</html>

Nothing here parses as a directive: no User-agent rule, no Sitemap. This is the status: no-directives branch, and the same applies to a file containing only comments, or only prose. The presence of the file is what hides the problem — the site looks configured and is not.

Corrected State (Passes): Fix the routing so /robots.txt is served as a plain-text file, with at least one real directive in it:

User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml

How PixyScan detects this

  1. Fetch. The crawler requests {origin}/robots.txt once per scan. This check only runs when the file came back with HTTP 200 — anything else is No robots.txt file on the site, not this issue.

  2. Content gate. A zero-length or whitespace-only body stops here and raises the missing issue instead. There is no configuration to correct, so calling it misconfigured would point you at content that does not exist.

  3. Line walk. The remaining body is walked line by line, the same way the parser walks it: # comments stripped, blank lines ending the current User-agent group. A line is recorded as an error if and only if the parser would discard it — no : separator, an unrecognised directive name, a valueless User-agent, a rule outside any group, a non-numeric Crawl-delay, or a Sitemap that is not an absolute http(s) URL. Each error names its line number.

  4. Classification. If the file yields no User-agent rule and no Sitemap directive at all, the finding is raised with status: 'no-directives'. Otherwise, if any line failed the walk, it is raised with status: 'syntax-errors'. A file that parses cleanly raises nothing.

  5. One row per scan. Both branches carry the full list in details.syntaxErrors and the count in details.syntaxErrorCount, and the message quotes the first error so the issue list is actionable without opening the detail page. A file that satisfies both branches produces a single finding.

  6. Scope. This is a site-level check — one finding per scan, not one per URL.

What we store

Storage Level

Site Level — one row per scan, keyed by scanId, not per URL. The finding is written with urlId NULL.


Database Table / Prisma Model

SiteCrawlBehaviourData — the parsed robots.txt payload.

AuditIssue — the finding itself, with the unusable lines in details.


Stored Fields

SiteCrawlBehaviourData

Field Type Description
robotsTxtData Json? The robots.txt payload. raw is the file byte for byte — original casing, comments, blank lines and directive order — which is what makes the line numbers in the finding resolvable. rawTruncated marks a file cut at 512k chars. userAgents, disallowedPaths, allowedPaths, crawlDelays and groups are the parsed rules; a line this issue reports is, by definition, absent from them. null when the file was not fetched or came back empty.

AuditIssue

Field Type Description
issueCode String robots_txt_misconfigured
scanId String The scan this finding belongs to
urlId String? NULL — the check is site-level
details Json? status (no-directives or syntax-errors), syntaxErrors (every unusable line, each prefixed with its line number), syntaxErrorCount, and message (quotes the first error so the issue list is actionable without opening the detail page)

The per-line list lives only in details. The crawler also computes a syntaxValidation block alongside robotsTxtData, but no DTO declares it, so the global validation pipe strips it before it reaches Prisma — details.syntaxErrors is the durable copy.


Detection Dependencies

  • The following data sources are required to evaluate this issue:
  • robots.txt — the raw text of {origin}/robots.txt, needed line by line rather than parsed, since the lines this issue reports are exactly the ones parsing discards
  • HTTP Response — the status code of that fetch; the check only runs on HTTP 200

Further reading