Skip to content

Product · Explore

robots.txt and the sitemap, read against each other

They are the two files a site uses to tell a crawler what to fetch, and they contradict each other constantly. A sitemap listing URLs robots.txt disallows is invisible if you can only look at one of them at a time — so this screen does not let you.

app.pixyscan.com/w/…/s/…/crawlability
The crawlability screen: robots.txt rules and sitemap coverage compared side by side.
Pinned in this screen’s header:Export robotsExport sitemap

What it answers

The questions this screen exists for

If none of these is a question you have, this is not the lens you want — one of the other ten probably is.

  • 01Is anything important being blocked by robots.txt?
  • 02Does the sitemap list URLs that no longer exist?
  • 03Are there pages on the site that the sitemap never mentions?
  • 04Which AI crawlers are allowed in?

On the screen

What you actually get

Four things, each of which is visible in the capture above.

  • The parsed robots.txt with its rules, beside the sitemap tree it is being read against.

  • Listed-not-crawled and crawled-not-listed, as two lists you can act on.

  • Sitemap lastmod recorded per URL, so a stale feed is visible rather than assumed.

  • Two exports, because this is the screen whose output people take to someone else.

Behind it

The checks that feed this screen

Every finding on this lens comes from a check in the catalogue — and every one of those carries a written guide to putting it right, rendered beside the finding in the product.

Browse the catalogue by discipline

The other lenses

One crawl, eleven readings of it

They are not separate scans. Each of these is the same dataset asked a different question.

See Crawlability against your own site

Connect one site free and run a crawl. Every lens on this page fills with your own pages within a few minutes.

Free plan · 1 site · 500 URLs a month · all 100+ checks · no card