Skip to content
Guide contents

User guide4 min

Crawlability

The Crawlability screen shows what search engines and AI crawlers can reach on your site, and what your site tells them. It covers robots.txt, sitemaps, redirects, each page's indexing rules and your language tags.

How to get there#

Crawlability is in the Explore group of the site sidebar, after Structure.

  1. 1

    Open Crawlability in the sidebar

    Open the site, then choose Crawlability under Explore.

  2. 2

    Start on Overview

    The Crawler journey card shows four steps in the order a crawler meets them: Crawler access (robots.txt), Discovery (sitemaps), Reach (how much of your sitemap the scan reached) and Redirect path (loops and long chains). Click a step to open its tab.

  3. 3

    Fix the steps from left to right

    A problem early on hides the ones after it. For example, if robots.txt blocks everything, the sitemap figures mean nothing.

app.pixyscan.com/w/…/s/…/crawlability

The Crawlability overview: the crawler journey - crawler access, discovery, reach, redirect path - above the What this site tells crawlers card, with six tabs and export buttons in the header.
The Overview tab. Below the journey, What this site tells crawlers sums up your robots.txt, sitemaps and crawl delay.

The six tabs#

A summary, then one tab for each thing a crawler reads.

FieldAnswersWhat it does
OverviewCan crawlers get around?The crawler journey, plus what your site tells crawlers: whether robots.txt loads, whether it names a sitemap, image and video sitemaps, and any Crawl-delay (a request to wait between fetches).
robots.txtWho is allowed in?Your robots.txt (the file that tells crawlers which paths to skip), rule by rule, grouped by user agent (the crawler's name). Below it, each known AI crawler is marked Allowed, Blocked, or Blocked by the wildcard rule.
SitemapsWhat did you list?Each sitemap file (your list of pages for search engines), how many URLs it lists, how many the scan reached, and how many have a lastmod (last changed) date.
RedirectsWhere do pages send you?Redirect chains, loops, chains over three hops and the longest chain. Split into redirects inside your site and pages that send visitors to other sites.
DirectivesWhat does each page say?Each page's canonical URL (the version you want in search), X-Robots-Tag and noindex (a request to stay out of search), whether it is indexable, and whether robots.txt allows it. Other views cover head tags, link tags, canonical faults and pagination faults.
InternationalWhich language, for whom?Your hreflang tags (which version of a page to show for each language or region), the x-default (the fallback version), whether each alternate links back, and any invalid language codes.

robots.txt

The robots.txt tab: groups of Allow and Disallow rules, then a grid of AI crawlers each marked Allowed or Blocked with its vendor and type.
The robots.txt tab shows your rules, then what they mean for each named crawler.

AI crawlers#

PixyScan knows 18 AI crawlers from 11 vendors. It reads your robots.txt the way each one would.

FieldWhat it doesWhat it does
retrievalCan cite youFetches pages so an AI search or answer tool can quote and link to your site.
on-demandFetches on requestFetches a page only when a user asks the AI tool about it.
trainingMay train modelsCollects pages that may be used to train AI models.

Choose which AI crawlers to allow

Many sites allow retrieval crawlers, so they can be cited, and block training crawlers. Decide what you want, change your robots.txt, then scan again and check AI readiness.

What to look for#

The most useful findings are where your sitemap, robots.txt and redirects disagree.

A sitemap page that robots.txt blocks sends search engines mixed signals. So does a sitemap full of URLs that redirect.

On the Sitemaps tab, compare “listed” with “reached” for each file. Many URLs never reached usually means the file lists old or blocked addresses.

A Crawl-delay of 10 seconds or more is flagged. It limits how many pages a crawler can fetch in a day, and Google ignores it anyway.

The header offers Export rules, Export sitemaps, Export redirects or Export hreflang, depending on the tab. Exports are included from Hobby up. To fix a problem, find the matching finding and its fix guide on Issues.

Does something here not match what you see in the app? Tell us