Skip to content
Issue docs

Soft 404: an error page returned with HTTP 200

Importantsoft_404_pageIssue 129

What is this issue?

A soft 404 is a page that tells the reader it is an error while telling the crawler it is a working page. The screen says "Page not found"; the server says HTTP 200 OK.

It happens by accident more often than by design. A framework catches an unknown route and renders an error component without setting a status. A CMS redirects every missing page to a "sorry" page that returns 200. A product is delisted and the template quietly falls back to an empty state.

For a page to pass this check, any one of these is enough:

  • Its headline is not an error statement; or
  • the server returned any status other than 200 — a real 404 or 410 is the right answer, and any other error status is reported by the broken-pages check; or
  • there is real content behind the headline — 200 words or more.

Example: /products/discontinued-item returning 200 OK with an <h1> of "Page not found" fails. The same page returning 404 Not Found passes — that is what a missing page is supposed to do.

An article about errors is not a soft 404. A post titled "How to fix a 404 error" is a real page about a real subject, and this check is written so that it is never reported.

Nor is a reference page whose headline is the status name. MDN's own "404 Not Found" page carries exactly the headline an error screen carries, and so does every API documentation site's status-code page. Nothing in the headline can tell them apart — what can is that they have several hundred words behind them, and an error screen has a sentence and a link home.

Why it matters

The status code is how a site tells a search engine what happened. Returning 200 for a missing page tells it the wrong thing, and the consequences are all bad:

  • Google drops the page from the index anyway. It detects soft 404s and files them under "Soft 404" in Search Console's Page Indexing report — not indexed. So the page is out of the index, and the site owner has no error in their logs to explain why.
  • Crawl budget is spent on nothing. A real 404 tells a crawler to stop coming back. A 200 tells it this is a live page worth recrawling, so it keeps returning to a page that will never have anything on it. On a large site with many dead routes, this is a substantial share of the crawl.
  • Links into it are invisible. A broken link that returns 404 shows up in every link report. One that returns 200 shows up in none of them, so the broken link is never found and never fixed.
  • The user experience is worse, not better. Browsers, bookmarking tools and link checkers all rely on the status. A friendly error page with a correct 404 is a better page, not a worse one.

It is graded IMPORTANT: the cost is a measurable indexing and crawl cost of exactly the kind that grade names, but the detection reads a headline rather than a status line, so it does not claim CRITICAL's mechanical certainty.

How to fix it

  1. Return the right status. A page that never existed, or that has been removed with nothing to replace it, returns 404. A page deliberately and permanently withdrawn returns 410. Both are correct; both tell a crawler to stop asking.

  2. Set the status where the error is decided, not in the template. The commonest cause is a route handler that renders an error component and never touches the response status. The error page and the status code have to be set together, or they drift apart again the next time someone edits the template.

  3. Redirect only when there is somewhere to go. If the content moved, a 301 to the new address is better than a 404. If it did not move, do not redirect to the homepage: that is a soft 404 with extra steps, and it is one of the patterns Google names explicitly.

  4. Keep the friendly error page. A helpful 404 — with search, with popular links, in the site's own design — is good practice. The status code is what changes; the page can stay exactly as it is.

  5. Check the empty states. A category with no products, a search with no results and a filter that matches nothing are legitimately 200 pages, but their headline should not read like an error. "No results for 'xyz'" is a working page; "Page not found" on the same URL is not.

  6. Watch the Soft 404 row in Search Console. It is the same finding from the other side, and it is the fastest way to confirm the fix landed.

Examples

Example 1: A framework catch-all route

Scenario: An unknown path falls through to a catch-all handler that renders the error component.

Fails because: the component is rendered, the status is never set, and the response goes out as 200.

GET /products/old-sku      →  200 OK
<title>Page not found</title>
<h1>Page not found</h1>

Corrected version:

GET /products/old-sku      →  404 Not Found
<title>Page not found | Acme</title>
<h1>Page not found</h1>

The page is identical. Only the status line changed, and that is the whole fix.

Example 2: An empty state that reads like an error

Scenario: A category page with no products left renders the site's shared "nothing here" component.

Fails because: the headline is an error statement, the status is 200 and there are eleven words on the page.

GET /shop/clogs           →  200 OK
<title>Nothing here</title>
<h1>Nothing here</h1>

Corrected version: either say what actually happened and keep the page useful, or return 404 if the category is gone for good.

<h1>No clogs in stock right now</h1>
<p>Try our <a href="/shop/boots">boots</a>, or leave your email…</p>

Note: redirecting every unmatched URL to the homepage is also a soft 404, and Google names that pattern explicitly — but PixyScan cannot see it here, because the homepage it lands on is a real page with a real headline. Return 404 instead.

Example 3: A page that is not a soft 404

Scenario: A blog post titled "How to fix a 404 error on your site", served with 200.

Passes, because the headline is not an error statement — it is an article about one. The check compares the whole headline rather than searching it for a phrase, precisely so that this page is never reported.

How PixyScan detects this

  1. Reads the status the server returned. Only a page that answered 200 is considered here — or one on a crawl path where the status was not recorded but the HTML arrived, since the page plainly served. Any other status means the server is already saying something, correctly (404, 410) or otherwise, and the broken-pages check is what reports the otherwise.

  2. Reads two headlines, and only two. The <title> and the first <h1>. Body text is deliberately not read: "page not found" appears in the footer, the search box and the help centre of plenty of working sites.

  3. Strips the site name. "Page not found | Acme Shop" is an error page, and "Page not found" is the part that says so. Everything from the first separator — a pipe, a dash, an en dash, a bullet, a spaced colon — is dropped, along with an opening "Oops!" or "Sorry,".

  4. Requires the whole headline to BE an error statement. This is the difference between a useful check and a noisy one. The headline is compared end to end against the statements error screens actually print — "404", "Page not found", "The page you are looking for does not exist", "We can't find that page", "Page unavailable" and the like. A headline that merely contains one of those phrases does not match, so an article about 404 errors is never reported.

  5. Requires the page to be empty as well. This is the condition that separates an error screen from a page about errors. An error screen is a headline, an apology and a link home; a status-code reference page has several hundred words behind the identical headline. A page with 200 words or more of body text is never reported, whatever its headline says.

  6. Steps aside for a noindexed page. A page carrying noindex is already out of the index, and the indexability check has already said so — reporting both would charge one page twice for one state.

  7. Reports what matched. The finding names the status code, whether the title or the H1 matched, the exact text that matched, and how many words were behind it.

Because a soft 404 is empty by construction, the thin-content check (#127) steps aside for any page reported here — one page that should simply return 404 is one row, not two.

What we store

Storage Level

Page Level


Database Table / Prisma Model

audit_issues.details

The inputs the check reads are already stored on PageSeoBasicsData (title, wordCount, indexabilityHttp, indexabilityNoindexAbsent); this check adds no column of its own.


Fields Used

Field Type Description
statusCode Int The status the response carried, or null if it was not read
matchedIn String title or h1 — which headline read as an error
matchedText String The headline text, exactly as the page wrote it
wordCount Int Words of body text behind the headline

Detection Dependencies

  • HTTP Response (status code)
  • HTML Document (title, first H1, and the body word count)
  • Robots directives (the meta robots tag and the X-Robots-Tag header, for the noindex suppression)

Note

The finding stores the headline verbatim rather than the normalised form it was matched against, so the row shows the reader what is actually on the page — which is what makes a false positive, if one ever occurs, obvious on sight rather than something to be inferred.

Further reading