Skip to content
Issue docs

Internal pages returning a server error (5xx)

Criticalinternal_server_error_pageIssue 165

What is this issue?

A page on your own site answered with a server error — an HTTP status in the 500 range — when PixyScan requested it. The page exists as far as your links are concerned, but the server failed to produce it.

The common ones:

Status What the server is saying
500 Internal Server Error The application crashed or threw while building the page
502 Bad Gateway A proxy or CDN in front of the site got a bad answer from the origin
503 Service Unavailable The server is overloaded, in maintenance, or rate limiting
504 Gateway Timeout The origin took too long to answer the proxy in front of it

For a page to pass this check:

  • Every internal page PixyScan fetched, or found linked from another page, answers without a 5xx status.

Example: /checkout is linked from every product page, and answers 500 because a database query fails for anonymous visitors. Readers following the link see an error page; search engines see a page that cannot be served.

How this differs from "Broken internal links and pages" (#8). #8 reports pages that answer 4xx — a page that was removed, renamed or never existed, which is a content or link fix. This check reports pages that answer 5xx — the page should exist and the server failed to produce it, which is a server or hosting fix, usually for a different person. One page is only ever reported under one of the two.

Why it matters

  • The page cannot be served to anyone. A reader who follows a link to it sees an error instead of the content, and leaves. If the page is a checkout, a sign-up or a pricing page, that is lost business, not just lost traffic.

  • Search engines back off from servers that fail. Google treats 5xx answers as a sign the server is struggling and slows down its crawling of the whole site. A page that keeps answering 5xx is eventually dropped from the index.

  • It often points at a wider problem. A single 500 is usually a bug on one route; many 5xx pages in one scan usually mean an overloaded origin, a misconfigured proxy or a deployment that is half-broken. Both are worth knowing before customers report them.

  • AI search and answer engines cannot read it either. An assistant fetching the page to answer a question about your product gets the error page, and answers without you.

Effect on the health score

This is a critical issue under the Link Integrity lens, the same grade as #8: the page cannot be served at all.

How to fix it

  1. Reproduce it. Open the reported URL in a private browser window (logged out, no cookies). Many 500s only happen for anonymous visitors or for a crawler's user agent.

  2. Read the server logs for that URL and time. The application log usually names the exception; the proxy or CDN log tells you whether the 502/504 came from the edge or the origin.

  3. Fix by status class:

    • 500 — an application error. Fix the code path, or make the page degrade gracefully when a dependency (database, API) fails.
    • 502 / 504 — the proxy cannot reach the origin, or the origin is too slow. Check origin health, upstream timeouts and keep-alive settings.
    • 503 — overloaded, in maintenance, or rate limiting. If it is maintenance, send a Retry-After header and keep the window short. If it is rate limiting, allow well-behaved crawlers through.
  4. If the page should not exist, remove the links to it and answer 404 or 410, or redirect it to the right page with a 301. A deliberate "gone" is better for readers and crawlers than a page that errors.

  5. Re-scan to confirm the page now answers 200.

When it fixes itself

A 503 during a short deployment or a traffic spike can disappear on the next scan. If the same page keeps appearing, it is not transient.

Examples

Example 1: A route that crashes for anonymous visitors

Fails: the product page links to the cart, and the cart throws when there is no session.

GET /cart
HTTP/1.1 500 Internal Server Error

Reported under this check against /cart, with statusCode: 500.

Passes: the route handles a missing session and renders an empty cart.

GET /cart
HTTP/1.1 200 OK

Example 2: A CDN that cannot reach the origin

Fails:

GET /blog/launch-notes
HTTP/1.1 502 Bad Gateway
server: cloudflare

The page exists at the origin, but the edge could not fetch it. The fix is in the origin's health or the proxy's configuration, not in any link.

Example 3: A removed page — reported elsewhere

GET /old-pricing
HTTP/1.1 404 Not Found

Not reported here. A 404 is a removed or renamed page and is reported under "Broken internal links and pages" (#8), which has a different fix: restore the page, redirect it, or update the links.

Example 4: A page that never answered — reported elsewhere

The request to /reports timed out after every retry. There is no status at all, so it is not a server error; it is reported as "Page failed to load or timed out during the crawl" (#193).

How PixyScan detects this

  1. When the crawler fetches a page itself. Every page PixyScan requests during the crawl has its HTTP status recorded. A page that answers 500–599 is recorded as failed with its real status, is not audited further (an error page has no title or headings worth checking), and is raised under this check. A page answering 400–499 is raised under #8 instead.

  2. After the crawl, for pages other pages link to. PixyScan reads the status of every internal link target. On plans that include link checks (Hobby and above), each distinct target is also requested once with a lightweight HEAD request (falling back to a one-byte GET if the server refuses HEAD), retried once on a transient 503. A target answering 5xx is raised here, once per page however many pages link to it.

  3. One code per page, decided in one place. The boundary — 500 to 599 is a server error, 400 to 499 is #8 — is the same rule at fetch time and after the crawl. If two passes saw the same page differently (for example the crawl's request got a 404 and a later probe got a 503), PixyScan keeps the finding that matches the page's most recent status, so the page never carries both.

  4. What is not reported. A request that never got an answer at all — a timeout, a DNS failure, a refused connection — has no status and is not a server error; it is reported as "Page failed to load or timed out during the crawl" (#193). A status above 599 is not a class a server can send and is left to #8.

  5. Respects your lens settings. Nothing is raised on a scan with the Link Integrity lens switched off.

What we store

Storage Level

Page Level


Database Table / Prisma Model

Url (urls) holds the status the finding is decided from. InternalLink (internal_links) is how a linked page is found after the crawl. The finding itself is stored on audit_issues.details.


Fields Used

Field Type Description
urls.status_code Int The HTTP status the page answered with. 500–599 raises this check; 400–499 raises #8. Null means never measured and is not judged
urls.status Enum failed for a page the crawler fetched and found answering an error
internal_links.target_num Int The linked page (by its urls.num), so every page some other page links to is judged once
audit_issues.issue_code String internal_server_error_page. Reconciled after the crawl so one page never carries this and no_broken_internal together

Stored Fields on the finding

Field Type Description
message String "Page returned a server error (HTTP 503)." or the post-crawl wording for a linked page
url String The page that answered with the server error
statusCode Int The 5xx status it answered with

Detection Dependencies

  • HTTP Response — the status of the page itself
  • Internal Links — which pages are linked, so linked pages are checked after the crawl
  • Plan — the post-crawl HEAD probe of link targets runs on Hobby and above; the fetch-time check runs on every plan

Further reading