Skip to content
Issue docs

Canonical tag outside the head

Importantcanonical_outside_headIssue 201

What is this issue?

A <link rel="canonical"> sits outside the document head — so search engines ignore it.

That happens in two ways:

  1. It is written in the body. A widget, a component or a CMS block outputs the tag inside <body>.
  2. The head closed before it. HTML parsers close the head the moment they meet an element that cannot be in it — a <div>, an <img>, an <iframe>, or stray text. Everything after that point is placed in the body, even when the source shows it inside <head>. A cookie banner or tracking pixel injected near the top of the head is the classic cause.

If the head has no other canonical, the page effectively declares none.

Why it matters

  • The canonical does nothing. Google reads rel="canonical" only from the head. A body canonical is ignored on purpose — the body is where user-generated content and third-party widgets live, and a canonical there could hijack the page.

  • It looks correct in the source. In the second case the tag is visibly inside <head> in the template; only the parsed document shows it in the body, so it survives code review.

  • Every other head tag after it is lost too. Whatever closed the head early also pushed any meta robots, hreflang or Open Graph tags that follow it into the body.

How to fix it

  1. Find the tag in the parsed document. Open the page in a browser's developer tools (Elements panel, not View Source) and look for link[rel=canonical]. If it is under <body>, something before it closed the head.

  2. If it is written in the body, move it into the template's head, or remove it if the head already has the right canonical.

  3. If the head closed early, move the offending element out of the head. Only title, base, link, meta, style, script, noscript and template belong there.

    <!-- Before: the <div> closes the head; the canonical lands in the body -->
    <head>
      <title>Red shoes</title>
      <div id="cookie-banner"></div>
      <link rel="canonical" href="https://example.com/shoes/red" />
    </head>
    
    <!-- After -->
    <head>
      <title>Red shoes</title>
      <link rel="canonical" href="https://example.com/shoes/red" />
    </head>
    <body>
      <div id="cookie-banner"></div>
    </body>
  4. Put the canonical early in the head, before any third-party scripts that inject markup.

Examples

Example 1: In the head

<head>
  <meta charset="utf-8" />
  <script src="/analytics.js"></script>
  <link rel="canonical" href="https://example.com/shoes/red" />
</head>

Passes.

Example 2: Optional tags omitted

<!doctype html>
<title>Red shoes</title>
<link rel="canonical" href="https://example.com/shoes/red" />
<h1>Red shoes</h1>

Passes. No <head> tag is written, but the parser places the canonical in the implied head.

Example 3: In the body

<body>
  <h1>Red shoes</h1>
  <link rel="canonical" href="https://example.com/shoes/red" />
</body>

Fails. The page effectively has no canonical.

Example 4: Head closed early

<head>
  <title>Red shoes</title>
  <img src="https://tracker.example/pixel.gif" />
  <link rel="canonical" href="https://example.com/shoes/red" />
</head>

Fails. The <img> closes the head; the canonical and everything after it are in the body.

Example 5: A stray copy in the body

<head><link rel="canonical" href="https://example.com/shoes/red" /></head>
<body><link rel="canonical" href="https://example.com/shoes" /></body>

Fails. The head canonical applies; the body tag is dead markup to remove.

How PixyScan detects this

  1. Every canonical link is found — rel read as a case-insensitive token list.

  2. Its position is decided the way an HTML5 parser decides it. A canonical with a <body> ancestor is outside the head. Otherwise, the document is walked in order up to the canonical: if anything appears that would have closed the head — an element other than title, base, link, meta, style, script, noscript or template, or non-blank text — the canonical is outside the head. Script, style, noscript and template contents are raw text to the parser and are not looked inside.

  3. Both crawl engines get the same answer. The browser engine hands over a DOM the browser has already rearranged; the default HTTP engine parses the HTML exactly as written, without moving anything. The walk above gives the same verdict for both, and a document that legitimately omits the optional <head> and <body> tags is not reported.

  4. Reported when any canonical is outside the head. The finding names the stray tag and, if there is one, the canonical in the head that actually applies.

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageSeoBasicsData (page_seo_basics_data). The finding itself is stored on audit_issues.details.


Fields Used

Field Type Description
canonical_url String? The first <link rel="canonical"> href in the document, exactly as written
has_duplicate_elements Boolean? True when the page declares a canonical (or title, description, H1) more than once
duplicate_elements Json? One entry per repeated element: { element: 'canonical', value, count }

Stored Fields on the finding

Field Type Description
canonicalUrl String The href of the first canonical outside the head
outsideHeadCount Int How many canonical tags are outside the head
headCanonicalUrl String? The canonical in the head, when there is one

Detection Dependencies

  • HTML Document — every canonical link and its position in the parsed document

Further reading