Skip to content
Issue docs

Removed pages returning 404 instead of a redirect

Importantmoved_deleted_returnIssue 20

What is this issue?

This issue checks whether pages that have been moved to a new URL or deleted return proper 301 redirects (for moved pages) or appropriate error pages (for deleted pages), rather than abrupt 404 errors. A 301 redirect passes link equity to the new URL, while abrupt 404s on moved pages lose accumulated PageRank.

A passing implementation means:

  • Pages that have moved to a new URL return a 301 (permanent) redirect to the new location
  • The redirect chain is short (ideally direct, not multiple hops)
  • Deleted pages that don't have a relevant replacement return 404 or 410 (Gone) status
  • No moved pages return 404 abruptly without a redirect

Example:

  • Old URL: https://example.com/old-page
  • New URL: https://example.com/new-page
  • ✅ Old URL returns 301 → https://example.com/new-page
  • ❌ Old URL returns 404 Not Found (link equity lost!)

Why it matters

Link Equity Preservation: 301 redirects pass the vast majority of link equity (PageRank) from the old URL to the new URL. Abrupt 404s lose this valuable ranking power.

User Experience: Users who bookmark or find old URLs in search results should be automatically taken to the new location, not encounter a 404 error.

Crawl Efficiency: Search engines will continue crawling old URLs that return 404, wasting crawl budget. 301 redirects tell them to update their index to the new URL.

SEO Rankings: Pages that have moved but don't redirect may lose their rankings because search engines see the content as "gone" rather than "moved."

Historical Value: Old URLs may have accumulated backlinks over time. 301 redirects ensure those links continue to benefit your site.

SEO Health Score: Proper redirect management is a fundamental technical SEO practice that significantly improves the site's health score.

How to fix it

  1. Identify moved pages - Review your site for pages that have been:

    • Moved to a new URL structure
    • Merged into other pages
    • Replaced with updated versions
  2. Implement 301 redirects - Set up server-side 301 redirects from old URLs to new URLs:

    • Apache: Redirect 301 /old-page https://example.com/new-page
    • Nginx: rewrite ^/old-page/?$ https://example.com/new-page permanent;
    • CMS: Use redirect plugins/modules
  3. Avoid redirect chains - Don't chain multiple redirects (A → B → C). Redirect directly to the final destination.

  4. Handle deleted pages - For pages that are deleted without a replacement:

    • Return 404 (Not Found) or 410 (Gone)
    • Consider a custom 404 page with helpful navigation
    • Don't redirect to the homepage (this confuses search engines)
  5. Update internal links - Change internal links to point directly to the new URLs instead of relying on redirects.

  6. Monitor with tools - Use Google Search Console to identify 404 errors and set up proper redirects.

Examples

Example 1: Proper 301 Redirect

Problematic State (Fails): Old page returns 404:

  • https://example.com/old-services returns 404 Not Found
  • Page has valuable backlinks that are now wasted

Corrected State (Passes): Implement 301 redirect:

  • https://example.com/old-services → 301 → https://example.com/services

Example 2: Redirect Chain

Problematic State (Fails): Multiple redirects in chain:

/old-page → /intermediate-page → /new-page

Corrected State (Passes): Redirect directly to final destination:

/old-page → /new-page

Example 3: Deleted Page Without Replacement

Problematic State (Fails): Page deleted but redirects to homepage (confusing for search engines):

  • https://example.com/discontinued-product → 301 → https://example.com/

Corrected State (Passes): Return 404 or 410:

  • https://example.com/discontinued-product returns 404 Not Found
  • Consider a custom 404 page with links to similar products

How PixyScan detects this

This check compares this scan against the previous one. That is the only honest way to ask it: a page can only be reported as "gone without a redirect" if PixyScan saw it when it was there.

1. Build the list of pages that have disappeared

At the end of the crawl, PixyScan asks the API for the pages that:

  • the most recent completed scan of this site on the same branch reached,
  • were actually fetched and answered 2xx in that scan — a URL left pending was never requested, and one that already answered 404 last time is not a page that has just gone away,
  • and that this crawl did not reach.

A URL this crawl reached is still there, so it is excluded before any request is made.

On a site's first scan this list is empty and the check reports nothing. There is no history to compare against yet, and that is the honest answer — this issue will begin reporting from the second scan onwards.

2. Probe each of them

Each remaining URL is requested with redirects not followed, a bounded number at a time, and the status code is read directly:

Response Verdict
301 or 308 Correct. A permanent redirect passes link equity to the new address. No finding.
302, 303, 307 Reported. A temporary redirect where a permanent one belongs.
404 or 410 Reported. The page is gone and left no redirect behind.
2xx No finding. A URL answering 200 has not moved — it simply was not linked from anywhere this crawl walked.
DNS failure, timeout, TLS error No finding. We could not reach it this time, which says nothing about whether it left a redirect. Recorded as unreachable in the scan data.

3. Attach each finding to its page

Every reported URL is resolved to its urls row through a bulk lookup, so each finding is a row about a specific address rather than one aggregate site-level note.

What this check does NOT do

  • It does not guess that a page "looks like" it moved from its URL shape or its content. Only a page previously seen and now unreachable is examined.
  • It does not report a healthy page. An earlier version fell back to the pages the crawl had just fetched — every one of which answers 200 — and reported the entire site as "moved pages not redirecting".
  • It does not measure redirect chain length. That is issue #10 (no_redirect_chains).

What we store

Storage Level

Page Level — This issue is evaluated for each individual page.


Database Table / Prisma Model

PageUrlParameterAudit


Stored Fields

Field Type Description
queryParamKeys String[] Array of query parameter keys found in the URL

Detection Dependencies

  • The following data sources are required to evaluate this issue:
  • URL — The crawler analyzes the URL structure and extracts query parameters
  • HTML Document — The page URL is parsed to identify parameters

Further reading