Skip to content
Issue docs

Malformed link URLs

Standardmalformed_link_hrefIssue 196

What is this issue?

A link on the page has an href that is not a usable web address.

Two kinds are reported:

  • An address the URL parser rejects outright. http:// with nothing after it, a space inside the host name (https://exa mple.com), a port number above 65535, an unclosed IPv6 bracket. A browser cannot follow these at all, and neither can a search engine.
  • A scheme typo the browser quietly "repairs" into a different address. https:example.com (no slashes) and http//example.com (no colon) do not fail — they are read as a path on YOUR site, so the link goes to https://yoursite.com/example.com or https://yoursite.com/blog/http//example.com, a page that does not exist.

For a page to pass this check:

  • Every <a href> is either a relative path, a complete http(s):// address, a fragment, or a valid non-web link such as mailto:, tel: or sms:.

Example: a footer link written as <a href="https:/www.partner.com"> looks right in the CMS, but visitors who click it land on a 404 page of your own site.

Why it matters

  • The link goes nowhere, or somewhere wrong. A visitor who clicks it gets an error or a 404 on your own site instead of the page you meant to send them to.

  • Search engines drop it. A crawler cannot follow an address it cannot parse, so whatever the link was meant to pass — discovery of the destination, the context of its anchor text — is lost.

  • The typo version creates junk URLs on your site. https:example.com is read as a path on your own domain. Crawlers request it, get a 404, and the site's error count goes up for a page that never existed.

  • It is invisible in the editor. The text of the link looks fine. Nothing tells the author the address is broken until somebody clicks it.

Effect on the health score

This is a standard-severity issue in the Link Integrity lens. It is raised once per page, however many malformed links the page carries.

How to fix it

  1. Find the link. Search the page source (or the CMS content) for the href shown in the finding. It is reported exactly as written.

  2. Write the full address. An external link needs the scheme, the colon and both slashes:

    <!-- Before -->
    <a href="https:/www.partner.com">Our partner</a>
    <a href="http//docs.example.com/guide">Guide</a>
    <a href="https://www.partner .com">Partner</a>
    
    <!-- After -->
    <a href="https://www.partner.com">Our partner</a>
    <a href="https://docs.example.com/guide">Guide</a>
    <a href="https://www.partner.com">Partner</a>
  3. Use a relative path for your own pages. href="/pricing" cannot be mistyped into a different domain.

  4. Fix it at the template if it repeats. A malformed link reported on every page is in a header, footer or sidebar template — fix it once there.

  5. Check generated links. An href assembled from a variable that was empty ("https://" + host) produces https://. Make the template skip the link when the value is missing.

Examples

Passes because: each href is a relative path, a complete address or a valid non-web link.

<a href="/pricing">Pricing</a>
<a href="https://docs.example.com/start">Docs</a>
<a href="//cdn.example.com/brochure.pdf">Brochure</a>
<a href="mailto:hello@example.com">Email us</a>
<a href="sms:+15555550100">Text us</a>

Example 2: A scheme typo

Fails because: https: with one slash is resolved as a path on this site.

<!-- On https://example.com/about -->
<a href="https:/partner.com">Partner</a>
<!-- A browser goes to https://example.com/partner.com -->

Corrected version:

<a href="https://partner.com">Partner</a>

Example 3: An address no browser can parse

Fails because: a host name cannot contain a space, and http:// alone has no host at all.

<a href="https://www.exa mple.com/shop">Shop</a>
<a href="http://">Website</a>

Corrected version:

<a href="https://www.example.com/shop">Shop</a>
<!-- and remove the empty one, or give it its address -->

How PixyScan detects this

  1. Reads every <a href> on the page while it extracts the page's links, from the HTML the server sent.

  2. Skips anchors that are not web links. Fragment-only links (#section), mailto: and tel: are skipped before any parsing. javascript: links are a different issue (#197). Other non-web schemes (sms:, geo:, whatsapp:) are valid by definition: the URL standard treats them as opaque, so they always parse and are never reported.

  3. Reports a scheme typo. An href starting http: or https: without the // that must follow, or http// / https// without the colon. A browser resolves both against the current page, so the link silently becomes a path on your own site.

  4. Reports an href the URL parser rejects. Each href is resolved against the page's address with the standard WHATWG URL parser — the same rules every browser uses. If it throws, the link has no address at all.

  5. Raises one finding per page with the number of malformed links and up to five of them exactly as they were written (whitespace trimmed, cut at 300 characters), so you can search the page source for them.

What we store

Storage Level

Page Level


Database Table / Prisma Model

audit_issues.details

Nothing else is stored: a malformed href has no address, so it cannot become a row in the link tables.


Stored Fields

Field Type Description
message String "Found N link(s) whose href is not a valid URL"
malformedLinkCount Int How many anchors on the page carry a malformed href
samples String[] Up to five of the hrefs, exactly as written (trimmed, cut at 300 characters)

Detection Dependencies

  • HTML Document — every <a href> on the page

Further reading