Removed pages returning 404 instead of a redirect
What is this issue?
This issue checks whether pages that have been moved to a new URL or deleted return proper 301 redirects (for moved pages) or appropriate error pages (for deleted pages), rather than abrupt 404 errors. A 301 redirect passes link equity to the new URL, while abrupt 404s on moved pages lose accumulated PageRank.
A passing implementation means:
- Pages that have moved to a new URL return a 301 (permanent) redirect to the new location
- The redirect chain is short (ideally direct, not multiple hops)
- Deleted pages that don't have a relevant replacement return 404 or 410 (Gone) status
- No moved pages return 404 abruptly without a redirect
Example:
- Old URL:
https://example.com/old-page - New URL:
https://example.com/new-page - ✅ Old URL returns
301 → https://example.com/new-page - ❌ Old URL returns
404 Not Found(link equity lost!)
Why it matters
Link Equity Preservation: 301 redirects pass the vast majority of link equity (PageRank) from the old URL to the new URL. Abrupt 404s lose this valuable ranking power.
User Experience: Users who bookmark or find old URLs in search results should be automatically taken to the new location, not encounter a 404 error.
Crawl Efficiency: Search engines will continue crawling old URLs that return 404, wasting crawl budget. 301 redirects tell them to update their index to the new URL.
SEO Rankings: Pages that have moved but don't redirect may lose their rankings because search engines see the content as "gone" rather than "moved."
Historical Value: Old URLs may have accumulated backlinks over time. 301 redirects ensure those links continue to benefit your site.
SEO Health Score: Proper redirect management is a fundamental technical SEO practice that significantly improves the site's health score.
How to fix it
Identify moved pages - Review your site for pages that have been:
- Moved to a new URL structure
- Merged into other pages
- Replaced with updated versions
Implement 301 redirects - Set up server-side 301 redirects from old URLs to new URLs:
- Apache:
Redirect 301 /old-page https://example.com/new-page - Nginx:
rewrite ^/old-page/?$ https://example.com/new-page permanent; - CMS: Use redirect plugins/modules
- Apache:
Avoid redirect chains - Don't chain multiple redirects (A → B → C). Redirect directly to the final destination.
Handle deleted pages - For pages that are deleted without a replacement:
- Return 404 (Not Found) or 410 (Gone)
- Consider a custom 404 page with helpful navigation
- Don't redirect to the homepage (this confuses search engines)
Update internal links - Change internal links to point directly to the new URLs instead of relying on redirects.
Monitor with tools - Use Google Search Console to identify 404 errors and set up proper redirects.
Examples
Example 1: Proper 301 Redirect
Problematic State (Fails): Old page returns 404:
https://example.com/old-servicesreturns 404 Not Found- Page has valuable backlinks that are now wasted
Corrected State (Passes): Implement 301 redirect:
https://example.com/old-services→ 301 →https://example.com/services
Example 2: Redirect Chain
Problematic State (Fails): Multiple redirects in chain:
/old-page → /intermediate-page → /new-pageCorrected State (Passes): Redirect directly to final destination:
/old-page → /new-pageExample 3: Deleted Page Without Replacement
Problematic State (Fails): Page deleted but redirects to homepage (confusing for search engines):
https://example.com/discontinued-product→ 301 →https://example.com/
Corrected State (Passes): Return 404 or 410:
https://example.com/discontinued-productreturns 404 Not Found- Consider a custom 404 page with links to similar products
How PixyScan detects this
This check compares this scan against the previous one. That is the only honest way to ask it: a page can only be reported as "gone without a redirect" if PixyScan saw it when it was there.
1. Build the list of pages that have disappeared
At the end of the crawl, PixyScan asks the API for the pages that:
- the most recent completed scan of this site on the same branch reached,
- were actually fetched and answered 2xx in that scan — a URL left pending was never requested, and one that already answered 404 last time is not a page that has just gone away,
- and that this crawl did not reach.
A URL this crawl reached is still there, so it is excluded before any request is made.
On a site's first scan this list is empty and the check reports nothing. There is no history to compare against yet, and that is the honest answer — this issue will begin reporting from the second scan onwards.
2. Probe each of them
Each remaining URL is requested with redirects not followed, a bounded number at a time, and the status code is read directly:
| Response | Verdict |
|---|---|
| 301 or 308 | Correct. A permanent redirect passes link equity to the new address. No finding. |
| 302, 303, 307 | Reported. A temporary redirect where a permanent one belongs. |
| 404 or 410 | Reported. The page is gone and left no redirect behind. |
| 2xx | No finding. A URL answering 200 has not moved — it simply was not linked from anywhere this crawl walked. |
| DNS failure, timeout, TLS error | No finding. We could not reach it this time, which says nothing about whether it left a redirect. Recorded as unreachable in the scan data. |
3. Attach each finding to its page
Every reported URL is resolved to its urls row through a bulk lookup, so each finding
is a row about a specific address rather than one aggregate site-level note.
What this check does NOT do
- It does not guess that a page "looks like" it moved from its URL shape or its content. Only a page previously seen and now unreachable is examined.
- It does not report a healthy page. An earlier version fell back to the pages the crawl had just fetched — every one of which answers 200 — and reported the entire site as "moved pages not redirecting".
- It does not measure redirect chain length. That is issue #10
(
no_redirect_chains).
What we store
Storage Level
Page Level — This issue is evaluated for each individual page.
Database Table / Prisma Model
PageUrlParameterAudit
Stored Fields
| Field | Type | Description |
|---|---|---|
| queryParamKeys | String[] | Array of query parameter keys found in the URL |
Detection Dependencies
- The following data sources are required to evaluate this issue:
- URL — The crawler analyzes the URL structure and extracts query parameters
- HTML Document — The page URL is parsed to identify parameters