Skip to content
Issue docs

Page more than three clicks from the home page

Suggestiondeep_page_click_depthIssue 143

What is this issue?

This is advice, not a defect. It reports a page that the crawl could only reach by following more than three links from the home page.

Click depth is a rough measure of how prominent a page is within its own site. The home page is depth 0; anything linked from it is depth 1; anything linked from those is depth 2, and so on. A page at depth 6 is one a reader has to work to find, and one that search engines will revisit less often than a page near the top.

For a page to pass this check:

  • It sits within three clicks of the home page along some route the crawl followed.

Example: /blog links to year archives, which link to month archives, which link to a paginated index, which links to the article. The article is depth 4, purely because of how the archive is built — and it is the only page of the four anyone wants to read.

What this number actually is. It is the depth the crawl reached the page at: the shortest route this crawl found, which is not provably the shortest route that exists. And it is not always recorded — pages the crawl reaches from the sitemap rather than by following a link have no depth at all. A page with no recorded depth is never reported here, and the report says how many of those there were rather than counting them as shallow.

That reliance on a proxy is a large part of why this is advice rather than a defect.

Why it matters

  • Crawlers reach it less often. Pages a long way from the entry points are discovered later and revisited less, so changes to them take longer to appear in search results.
  • It gets less of the site's authority. Internal links carry signal, and each hop divides what arrives. A page four or five levels down receives very little of what the home page has to give.
  • Readers do not find it. Nobody clicks four times looking for something. If the page matters, the route to it is worth shortening.
  • It usually reveals the structure. Depth is rarely chosen. It is the by-product of a date archive, an unpaginated category, or a menu that stops at the second level, and looking at the deep pages is a fast way to see it.

It is reported as a suggestion and carries no severity, so it never deducts from the health score. Two reasons, and both matter:

  1. A deep catalogue is a legitimate shape. A shop with 40,000 products, or an archive spanning fifteen years, genuinely has pages a long way down, and nothing about that is wrong.
  2. The evidence is a proxy. What is measured is where the crawl arrived, not a proven shortest path, and the measurement is missing on many pages. A finding resting on a proxy should not be deducting from a score — the same reasoning the catalogue applies to "Crawled pages are missing from llms.txt".

How to fix it

  1. Decide whether the page matters. Depth is only a problem for pages you want found. A 2011 archive page being five clicks down is fine.

  2. Link to it from higher up. The direct fix: add a link from a section index, a related-content block, or the navigation. One link from a depth-1 page makes it depth 2.

  3. Flatten date archives. Year → month → day → page is four levels before any content. Link articles from the section index instead and leave the date archive as a secondary route.

  4. Paginate sensibly. A category paginated 60 pages deep puts its last items an enormous distance from the top. Larger pages, sub-categories, or a filterable index all shorten it.

  5. Publish hub pages. A page that gathers the twenty best pieces on a topic pulls all twenty closer to the surface at once, and is worth reading in itself.

  6. Check the menu depth. A navigation that stops at the second level leaves everything below it depending on in-page links to be reachable at all.

Examples

Example 1: A shallow route to the same article

Scenario: The section index links articles directly.

Passes because: the article is two clicks from the home page.

/                       depth 0
/guides                 depth 1
/guides/technical-seo   depth 2

Example 2: A date archive standing between the reader and the content

Scenario: Articles are only linked from month archives, which are only linked from year archives.

Fails because: the article is four clicks down, entirely because of the archive structure.

/                          depth 0
/blog                      depth 1
/blog/2026                 depth 2
/blog/2026/03              depth 3
/blog/2026/03/seo-basics   depth 4   <- reported

Corrected version: link articles from /blog itself and keep the date archive as an alternative route.

/                          depth 0
/blog                      depth 1
/blog/2026/03/seo-basics   depth 2

Example 3: Deep pagination

Scenario: A category of 1,800 products paginated 30 at a time, each page linking only to the next.

Fails because: the products on page 40 are more than forty clicks down.

/shop/shoes            depth 1
/shop/shoes?page=2     depth 2
/shop/shoes?page=3     depth 3
/shop/shoes?page=40    depth 40   <- and every product on it

Corrected version: link the first, last and a spread of intermediate pages from every pagination bar, and add sub-categories so products are reachable by type rather than only by position in a list.

How PixyScan detects this

  1. Counts hops as the crawl walks the site. The starting address is depth 0. Every page discovered by following a link from a page at depth n is recorded at depth n + 1.

  2. Takes the first arrival. Pages are fetched in the order they were queued, so the first time a page is reached is by the shortest route the crawl found.

  3. Reports anything deeper than three. Depth 3 passes; depth 4 and beyond are reported, with the actual number on the finding.

  4. Says nothing about pages with no recorded depth. A page reached from the sitemap rather than by following a link has no hop count, and no claim is made about it in either direction. How many such pages there were is reported alongside, so a quiet result is never mistaken for a clean one.

  5. Ignores addresses that are not pages. A URL that redirects is a signpost rather than a page — telling somebody to shorten the path to it would be advice about nothing — and a page that answered 4xx or 5xx is already reported by the broken-pages check. Neither is included here.

  6. Reports the page itself, not whatever links to it, because it is the page that needs a shorter route.

What we store

Storage Level

Page Level


Database Table / Prisma Model

Url. The finding itself is stored on audit_issues.details.


Fields Used

Field Type Description
urls.crawl_depth Int Hops from the starting address along the route the crawl followed. NULL means this scan did not measure it, which is not the same as zero
urls.url String The page being reported
urls.status Enum Rows the URL quota refused or the exclude patterns dropped are not pages of the report and are left out
urls.status_code Int Only a page that served content (200-299) is judged; a redirect is a signpost and an error page belongs to the broken-pages check

Stored Fields on the finding

Field Type Description
clickDepth Int How many hops the crawl needed to reach this page
limit Int The threshold it crossed, stored on the finding so the row explains itself
url String The page being reported

Detection Dependencies

  • Internal Links — depth is accumulated as the crawl follows them
  • HTTP Response — the status that decides whether the row is a page at all

Further reading