Skip to content
Issue docs

Page failed to load or timed out during the crawl

Importantpage_failed_to_loadIssue 193

What is this issue?

PixyScan asked your server for this page, retried, and never got an answer: no status code, no headers, no HTML.

This is different from a page that answers with an error. A 404 or a 500 is a response, and other checks report it. This check reports pages where the request itself failed:

What happened Shown as
The server did not answer in time timeout
The domain name could not be resolved dns
The server refused the connection connection_refused
The connection was dropped before a response connection_reset
The HTTPS connection could not be set up (certificate, protocol) tls
Anything else that stopped a response arriving other

A page that redirects in a loop also fails to load. It is recorded with the class too_many_redirects and reported by the Redirect loop check instead of this one, so the same fault is not reported twice.

For a page to pass this check:

  • It returns a response, of any status, within the crawl's time limit.

Example: a product page runs a slow database query and takes a minute to respond. The crawler gives up at its time limit (45 seconds by default) on each attempt, and the page is reported as a timeout.

Why it matters

Search engines hit the same wall. Googlebot has time limits too, and it does not index a page it cannot fetch. If it keeps failing, it slows down its crawling of your whole site.

Visitors hit it first. A page that times out for a crawler usually times out, or nearly does, for people on slow connections.

It hides everything else. No other check can run on a page that never loaded, so the page has no title, links or content to audit. That makes it a blind spot in the report until it loads.

Why IMPORTANT. A failure here is a real fault more often than not. But a single timeout can be a passing network problem, so it is graded below a page that is definitely broken. If the same page fails scan after scan, treat it as broken.

How to fix it

Open the page in the issue list. The finding names the error class and the exact error message. Fix according to the class:

  1. timeout. Find out why the page is slow. Check server logs for the URL, slow database queries and blocking calls to external services. Add caching for pages that are expensive to build. Aim for the HTML to arrive in well under a few seconds.

  2. dns. Make sure the domain has A/AAAA records at the nameservers it is delegated to. If only some pages fail, check whether they are on a different subdomain.

  3. connection_refused / connection_reset. The server, a load balancer or a firewall is dropping connections. Check that the service is up and listening on 80/443, and that rate limiting or bot protection is not cutting off crawlers. Allow Googlebot and your monitoring tools if needed.

  4. tls. Renew an expired certificate, make sure it covers this hostname, and serve the full chain, including intermediate certificates. An SSL checker shows which part is wrong.

  5. other. Read the message in the finding, then reproduce the request from outside your network.

(A redirect loop is reported by the Redirect loop check, which explains how to break it.)

Then re-scan. A page that fails once and loads on the next scan was a temporary problem. A page that fails every time needs fixing.

Examples

Example 1: a timeout

$ curl -o /dev/null -s -w "%{time_total}\n" https://example.com/reports/annual
62.408

The page takes a minute. Each attempt is abandoned at the crawl's time limit (45 seconds by default).

Finding: fetchErrorClass: "timeout", message request timed out after 45 seconds.

Corrected: cache the report and serve it in under a second.


Example 2: an expired certificate

Finding: fetchErrorClass: "tls", message certificate has expired.

Corrected: renew the certificate (or fix auto-renewal) and serve the full chain.


Example 3: not reported here

  • A page answers 503 Service Unavailable quickly. That is a response, so the server-error check reports it.
  • A page redirects in a loop (/shop to /shop/ and back). It is stored with fetchErrorClass: "too_many_redirects" and reported by the Redirect loop check.

How PixyScan detects this

  1. Every response is handled as a response. A 4xx or 5xx page goes through the normal checks and is reported by the status-code checks, not here.

  2. A request that produces no response is retried. Only when every retry fails is the page treated as failed. The finding says how many attempts were made.

  3. The error is classified from the error code and message that the HTTP client or browser returned:

    • timeout: the request or navigation timed out (ETIMEDOUT, Timeout … exceeded, net::ERR_TIMED_OUT)
    • dns: ENOTFOUND, EAI_AGAIN, net::ERR_NAME_NOT_RESOLVED
    • connection_refused: ECONNREFUSED
    • connection_reset: ECONNRESET, socket hang up, net::ERR_EMPTY_RESPONSE
    • tls: certificate and SSL/TLS errors (CERT_HAS_EXPIRED, ERR_SSL_PROTOCOL_ERROR)
    • too_many_redirects: the redirect limit was hit (Redirected 10 times, ERR_TOO_MANY_REDIRECTS); see below
    • other: anything not recognised
  4. The page is recorded as failed, with status code 0, the class, and the error message (one line, at most 500 characters), and this issue is raised on it.

What is deliberately not reported here:

  • A page that loaded but whose processing by PixyScan ran out of time (Crawlee's "requestHandler timed out"). The page answered; the delay was on our side. It is still recorded as failed, with the class crawler, but no finding is raised against your site.
  • A redirect loop (too_many_redirects). The Redirect loop check reads that class and reports the page there. Reporting it here too would count one fault twice.

What we store

Storage Level

Page Level: raised on the page that could not be loaded.


Database Table / Prisma Model

Url, and AuditIssue (url_id set)


Fields Used

Field Type Description
Url.status Enum FAILED
Url.statusCode Int? 0: no HTTP response was received
Url.fetchErrorClass String? timeout, dns, connection_refused, connection_reset, tls, too_many_redirects, crawler or other
Url.fetchErrorMessage String? The error's own text, one line, at most 500 characters

fetchErrorClass and fetchErrorMessage are null on every page that loaded, and on failed pages from scans taken before these columns existed ("not recorded"). Pages with the class crawler or too_many_redirects have the columns set but no #193 finding: the first was a delay on PixyScan's side, and the second is reported by the Redirect loop check (#166), which reads this column.

Finding Details

Key Description
fetchErrorClass As stored on the page
fetchErrorMessage As stored on the page
retries Retries made after the first attempt

Detection Dependencies

  • HTTP request: the crawler's fetch of the page, after every retry
  • Error code and message: from the HTTP client (got) or the browser (Playwright)

Further reading