Skip to content
Issue docs

External link redirects

Suggestionexternal_link_redirectsIssue 203

What is this issue?

This is advice, not a defect. A link on your page points at an address on another site that redirects before it answers:

https://vendor.com/old-docs  →  301  →  https://vendor.com/docs/

The link works — the reader ends up on the right page — but every click pays for an extra round trip, and the link now depends on the other site keeping that redirect forever.

For a page to pass this check:

  • Every external link points at the address that answers, not at one that redirects to it.

Example: you linked to a partner's pricing page two years ago. They restructured their site and redirected the old address. Your link still works today; updating it to the new address keeps it working when they tidy up their redirects next year.

http:// links that upgrade to https:// are included. It is the most common case and the cheapest fix — one character — and it is marked as such on the finding.

Redirects that end on an error are not this check. If the redirect leads to a page that answers 4xx or 5xx, the link is broken and is reported as "External link returns a broken or error HTTP status code" (#97).

Why it matters

  • Every click waits longer. Each redirect is another request and response before the page starts loading — often a new connection to a new host, too.

  • The link outlives the redirect only by luck. Sites restructure and prune old redirects. A link that points at the final address keeps working; one that relies on a redirect breaks the day the other site removes it.

  • An http:// hop leaks the first request. Before the upgrade to HTTPS, the request — and the address of the page the reader came from — travels unencrypted.

  • Shorteners and tracking redirects hide where a link goes. Readers, and crawlers judging the link, see the shortener rather than the destination.

Effect on the health score

None, by design. This is a suggestion, so it carries no severity and deducts nothing.

How to fix it

  1. Open the finding and copy the final address for each link.

  2. Replace the link with the final address.

    <!-- Before -->
    <a href="http://vendor.com/old-docs">Vendor documentation</a>
    
    <!-- After -->
    <a href="https://vendor.com/docs/">Vendor documentation</a>
  3. Start with https-upgrade links. They are a one-character change and remove an unencrypted request.

  4. Check moved links make sense. If the final address is a home page or a login page rather than the content you meant, the content has probably gone — find its new home or remove the link.

  5. Leave deliberate redirects alone. Affiliate and tracking links redirect by design; if you need them, keep them.

When to leave it alone

  • Affiliate, tracking or shortened links you are required to use.
  • Links where the other site asks you to use a stable redirecting address.

Examples

Reported (kind: https-upgrade):

<a href="http://vendor.com/docs">Docs</a>
http://vendor.com/docs → 301 → https://vendor.com/docs   (1 hop)

Fixed:

<a href="https://vendor.com/docs">Docs</a>

Example 2: A page that moved

Reported (kind: moved, 2 hops):

https://partner.org/pricing-2023 → 301 → https://partner.org/plans → 301 → https://partner.org/plans/

Example 3: A redirect into a dead page

https://vendor.com/old → 301 → https://vendor.com/gone → 404

Not reported here. The link is broken and is reported under #97.

https://news.example.org/a → 302 → https://consent.example.org/?next=a → 302 → https://news.example.org/a → 200

Not reported. The link already points at the address that answers.

How PixyScan detects this

  1. Checks each distinct external address once. After the crawl, on plans that include link checks (Hobby and above), PixyScan requests every distinct external address with a lightweight HEAD request (falling back to a one-byte GET if the server refuses HEAD). This is the same request that finds broken external links; no extra request is made for this check, and it uses no extra link-check allowance.

  2. Records the redirects it followed. Up to five redirects are followed, as before. PixyScan now keeps how many it followed and the address that finally answered, on each external link row.

  3. Reports a redirect that ends on a working page. An address that redirected at least once and finally answered 2xx is reported, once per page that links to it, with the address it ends up at.

  4. What is not reported.

    • A redirect that ends on a 4xx or 5xx — that link is broken (#97).
    • A redirect that ends on another redirect PixyScan could not follow, or ran past five hops — there is no final address to recommend.
    • A "bounce" that ends on the very address the link already uses, such as a cookie or consent round trip (/article → /consent → /article) — there is nothing to change.
    • Addresses on localhost or private networks, which are never requested (see #168).
  5. Says what kind of redirect it was. Each reported link carries a kind: https-upgrade (only http: → https:), www (only the www. prefix changed), trailing-slash (only a trailing slash changed), or moved (anything else).

  6. One finding per page. Each page gets one finding listing up to 20 of its redirecting links, with the full count.

Known false positives. A site that redirects automated clients to a login or consent page that people never see will be reported. The finding names the final address, so such a row explains itself.

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageExternalLink (page_external_links). The finding is stored on audit_issues.details.


Fields Used

Field Type Description
page_external_links.external_url String The address the link points at
page_external_links.status_code Int The status of the address that finally answered. Only 200–299 is reported
page_external_links.redirect_count Int Redirects followed. 0 = answered directly; null = not checked (plan, budget, or the request failed)
page_external_links.final_url String The address that answered after the redirects. Null when nothing redirected or not checked
page_external_links.url_id String The page carrying the link, and the page the finding is raised against

Stored Fields on the finding

Field Type Description
message String How many redirecting external links the page carries
recommendation String What to change
redirectedLinks Array Up to 20 of { externalUrl, finalUrl, redirectCount, kind }, sorted by address
redirectedLinkCount Int How many distinct redirecting addresses the page links to
truncated Boolean True when there were more than 20

Detection Dependencies

  • External Links — every link to another host recorded by the crawl
  • HTTP Response — the redirects followed by the external link check
  • Plan — external link checking runs on Hobby and above

Further reading