Skip to content
Issue docs

External link returns a broken or error HTTP status code

Standardbroken_external_linkIssue 97

What is this issue?

An external link is any <a href> on your page that points at a different domain. This issue checks whether those outbound links still work — that is, whether the destination answers with a successful HTTP status code rather than an error one.

The web outside your site changes without telling you. A page you linked to two years ago gets deleted, a company rebrands and drops its old URLs, a documentation site restructures its paths. Your HTML still carries the link, so the page looks fine to you, and the reader is the one who finds out it is dead.

A passing implementation means:

  • Every outbound link on the page resolves to a successful status (2xx), or to a redirect that ends at one
  • No outbound link returns a client error (4xx) such as 404 Not Found or 410 Gone
  • No outbound link returns a server error (5xx) such as 500 Internal Server Error or 503 Service Unavailable
  • No outbound link returns an access refusal (401 Unauthorized or 403 Forbidden)

Example: A blog post cites a research paper at https://example-university.edu/papers/seo-study.pdf. The university reorganised its site and that path now returns 404. The citation still renders as a normal link in your article, but every reader who clicks it lands on an error page.

The issue is raised once per page and lists every broken destination found on that page — so a page carrying six dead links is one row to work through, not six.

What this issue does not report

  • Links that could not be reached at all. A timeout, a DNS failure, a TLS error or a refused connection means PixyScan has no measurement, so it asserts nothing. These are counted and explained in the scan log, but they are never reported as broken links.
  • Links skipped because the plan does not cover outbound checks. An unchecked link is left unchecked, not assumed broken.
  • Internal links. Same-domain links are covered by their own issue.

Why it matters

User Experience: A broken outbound link is a dead end in the middle of your content. The reader clicked because you told them there was something worth seeing, and the destination gave them an error page instead. On pages that exist to cite sources — documentation, research write-ups, comparison articles, resource lists — enough dead links make the whole page feel abandoned.

Content Quality Signals: Search engines assess whether a page is maintained. A page whose outbound references have rotted looks stale next to a competitor's page covering the same topic with links that still resolve, and staleness is exactly what quality-focused ranking systems are designed to notice.

Wasted Link Equity: Outbound links pass authority to the pages you cite. A link to a 404 passes it into a void. It is context you meant to give search engines about your topic, spent on a destination that no longer supports the claim.

Crawl Efficiency: Crawlers follow outbound links too. Sending them repeatedly to dead destinations spends crawl effort on nothing, and does it on every page carrying the same broken link.

AI Search / AEO: Answer engines lean on citations to judge whether a page is grounded. A page whose sources cannot be fetched is harder to trust and harder to quote, so it is less likely to be used as a source in a generated answer.

Trust and Credibility: For commercial pages, broken links to partners, certifications, reviews or payment providers directly undermine the case the page is making.

SEO Health Score: This issue is scored as STANDARD severity at page level. Because outbound links tend to be repeated in shared components — a footer, a sidebar, an author bio — a single dead destination can raise the issue on every page of the site at once. Fixing it once in the template clears the finding site-wide and produces a disproportionate improvement to the technical health score.

How to fix it

  1. Open the page report and read the status code for each broken destination. The code tells you which fix applies, and they are not interchangeable:

    • 404 / 410 — the page is gone. Find the replacement or remove the link.
    • 500 / 502 / 503 — the destination is failing right now. Recheck before acting; this is often temporary.
    • 401 / 403 — the destination refuses automated clients, or the content was moved behind a login. Open it in a browser to tell the two apart.
  2. Repoint the link to the live equivalent first. Deleting a citation loses the context it gave the reader. Check the destination site's own search, or an archive of the old URL, for where the content moved. Updating the href is almost always better than removing it.

  3. If nothing replaces it, remove the link but keep the text. Turn <a href="https://gone.example/report">the 2023 industry report</a> into plain text, or cite an archived copy. The claim stays; the dead end goes.

  4. Fix it in the template, not on each page. If the broken link appears on dozens of pages, it lives in a header, footer, navigation block or reusable component. Change it there once.

  5. Verify a 403 before you touch it. Many large sites answer automated requests with 401 or 403 while serving the same URL perfectly to a browser. If the link works when you open it, the destination is fine and no edit is needed — the status-code filter on the Links screen is how you take a code you know is bot protection back out of the list.

  6. Point links at the final destination. If a link redirects several times before landing, update the href to the end of the chain. Redirect chains are where links break in the first place — every hop is another thing that can be removed later.

  7. Prefer stable, canonical URLs when adding new links. Link to a documentation root rather than a versioned deep path; avoid URLs carrying session identifiers or campaign parameters; use https where the destination supports it.

  8. Recheck on a schedule. Outbound links break through no action of yours, so a page that passes today can fail next month. Run the scan periodically rather than treating this as a one-time cleanup.

Examples

Example 1: A cited source that was deleted

An article links to a study that the publisher has since removed.

Problematic State (Fails):

<p>
  According to the
  <a href="https://research.example.org/2021/mobile-speed-study">
    2021 mobile speed study
  </a>, most users abandon a page after three seconds.
</p>

https://research.example.org/2021/mobile-speed-study → 404 Not Found

The claim is still on the page, but the evidence behind it is a dead end.

Corrected State (Passes):

<p>
  According to the
  <a href="https://research.example.org/reports/mobile-speed-2021">
    2021 mobile speed study
  </a>, most users abandon a page after three seconds.
</p>

https://research.example.org/reports/mobile-speed-2021 → 200 OK

The publisher moved the report under /reports/. Repointing the link keeps the citation intact. If no replacement existed, the better fix would be to drop the <a> and leave the text — the claim survives, the dead end does not.


A partner logo in the site footer points at a page that no longer exists after the partner's rebrand.

Problematic State (Fails):

<!-- footer.html — included on all 240 pages -->
<a href="https://oldpartner.example.com/certified-agencies">
  <img src="/img/partner-badge.png" alt="Certified Partner" />
</a>

https://oldpartner.example.com/certified-agencies → 410 Gone

PixyScan makes one request to that destination and then raises the issue on all 240 pages, because all 240 carry the link. The issue list shows 240 affected pages for what is a single edit.

Corrected State (Passes):

<!-- footer.html -->
<a href="https://newpartner.example.com/partners/certified">
  <img src="/img/partner-badge.png" alt="Certified Partner" />
</a>

https://newpartner.example.com/partners/certified → 200 OK

Fixing the shared template once clears the finding from every page. This is why it is worth sorting the affected pages by what they have in common before editing anything.


Example 3: A 403 that is not actually broken

A link to a large news site is reported as broken, but opens perfectly in a browser.

Reported State:

<a href="https://news.example.com/business/seo-market-report">
  industry coverage
</a>

https://news.example.com/business/seo-market-report → 403 Forbidden

What is happening: the destination's bot-protection layer refuses the automated request while serving the same URL normally to a real browser. PixyScan reports the response it actually received rather than guessing at intent — the alternative, suppressing all 403 responses, would mean a genuinely locked-down page could never be reported at all.

How to resolve it:

  1. Open the URL in a browser. If it loads, the link is fine and needs no edit.
  2. Use the status-code filter on the Links screen to take 403 out of your working list once you have confirmed it is your own or the destination's WAF.
  3. If the browser also shows an error or a login wall, the content really has moved behind access control — treat it as Example 1 and repoint or remove the link.

How PixyScan detects this

This check runs after the crawl has finished, once every page's outbound links have been collected. It works in five logical steps.

  1. Collect outbound links while crawling. On every page, PixyScan reads each <a href>, resolves it against the page URL, and compares its host with the site's own host. Anything on a different host is an external link and is recorded together with the anchor text a reader would see. Anchors that are not requests at all — a bare # fragment, a mailto: address, a tel: number — are skipped, as is any href that is not a valid URL.

  2. Reduce the links to distinct destinations. All recorded links across the whole scan are grouped by destination URL. A social link repeated in the footer of a 75-page site is one destination, not 75 links. This matters for accuracy as much as for speed: contacting one host 75 times in a few seconds is what rate limiting exists to stop, and the throttled responses that came back would be failures PixyScan had manufactured itself.

  3. Request each distinct destination once. PixyScan sends a lightweight HEAD request, which asks for the response status without downloading the page body, and follows up to five redirects to reach the final destination. If the server refuses HEAD, the request is retried as a GET. Each attempt has a ten-second deadline, and a request that fails is retried once before being given up on.

  4. Fan the answer back out to every link. The status measured for a destination is applied to every link pointing at it, so two pages linking to the same URL can never disagree about whether it works.

  5. Judge the status and raise the issue. A destination is broken when the server itself said so — any status of 400 or above, which includes 401, 403 and 429. The issue is raised once per source page, carrying every broken destination found on that page along with its status code.

When the check does not run, or does not report

  • The destination could not be reached. A timeout, DNS failure, TLS error, refused connection or redirect loop produces a reason, not a status. Nothing is stored and no issue is raised, because "we could not reach it" is not something PixyScan can tell you is broken. Any status recorded by an earlier successful scan is left untouched. These are counted, and the reasons are reported in the scan summary.
  • The workspace plan does not cover outbound requests. No destination is contacted, nothing is recorded, and every link stays marked as unchecked rather than assumed broken.
  • The monthly outbound-request allowance runs out part-way through. The destinations not yet reached are left unchecked for the same reason.
  • The Crawl Behaviour toggle is switched off for the site. The check does not run at all.

Known limits

  • Only links present in the HTML the server returned are checked. Links a page builds with JavaScript after load are not seen.
  • A 401 or 403 is reported as broken. Many large sites answer automated clients that way while serving the same URL normally to a browser. PixyScan reports the response it received and leaves that judgement to the reader, who can filter those codes out on the Links screen.
  • A destination is measured once per scan. A site that happened to be down during the scan window is reported as broken for that scan and clears on the next one.

What we store

Storage Level

Page Level — the finding is written against the source page that carries the broken link, keyed by urlId. One finding per page, however many of its outbound links are broken.

The link measurements themselves are stored per link occurrence, but the status is resolved once per distinct destination and applied to every occurrence of it in the scan.


Database Table / Prisma Model

PageExternalLink — one row per outbound <a> occurrence, carrying the measured status code.

AuditIssue — the finding itself, with every broken destination on the page listed in details.


Fields Used

Field Type Description
id String Identifier for this link occurrence
urlId String The source page carrying the link — what the finding is attributed to
scanId String The scan this measurement belongs to
externalUrl String The absolute destination URL, resolved against the page URL. Also the grouping key: every row sharing a value is answered by one request
statusCode Int? The HTTP status the destination returned. NULL means not measured — never probed, skipped by the plan, or unreachable. A status is written only when a real one was received; no placeholder is ever stored
anchorText String? The visible link text, falling back to an image alt for image links and capped at 300 characters, so a reported link can be located on the page

AuditIssue

Field Type Description
scanId String The scan this finding belongs to
urlId String? The source page carrying the broken link or links
issueCode String broken_external_link
details Json? brokenLinks (each entry an externalUrl and the statusCode it returned), brokenLinksCount, and message summarising how many were found on this page

A uniqueness constraint on (scanId, urlId, issueCode) is what makes this one finding per page: all of a page's broken destinations are collected into details rather than written as separate rows.


Detection Dependencies

  • HTML Document — every <a href> on the crawled page, from which outbound links are identified by comparing host against the site's own
  • HTTP Response — the status code returned by each distinct destination, obtained by a post-crawl HEAD request, falling back to GET
  • Redirect Chain — redirects are followed to the final destination, up to five hops, and the status is judged there
  • Plan Entitlement and Outbound-Request Allowance — determine whether destinations are contacted at all; without them the links are recorded as unchecked

Source of Truth

This repository does not currently contain a Database Schema Mapping.md. The fields above were taken from the Prisma schema (apps/api/prisma/schema/seo-links.prisma and apps/api/prisma/schema/audit-issues.prisma) and from the pass that writes them. If the mapping document is added later, it takes precedence over this file.

Further reading