Skip to content
Issue docs

hreflang pointing at a page search engines cannot use

Importanthreflang_target_not_indexableIssue 145

What is this issue?

An hreflang tag is an instruction to index a particular address as the translation of this page. The instruction only works if the address is one a search engine can actually index.

This check reports a page whose hreflang set points at an address that

  • does not answer HTTP 200 — it 404s, it 500s, or it redirects somewhere else, or
  • answers 200 but carries a noindex directive.

For a page to pass this check:

  • Every href in its hreflang set resolves to a live, indexable page.

Example: a site retires /de/preise and moves it to /de/tarife, adding a 301. The redirect is correct, and every internal link follows it. The hreflang blocks were not updated, so every other language version still names /de/preise. Google follows the annotation, finds a redirect rather than a page, and drops German from the cluster.

Why it matters

A dead alternate is a missing translation, as far as search is concerned. Google will not substitute the redirect's destination for the address you declared. The pair is dropped, and readers in that locale are shown whichever version ranks on its own merits.

Redirects are the commonest form of this, and the least visible. A 301 is a healthy thing everywhere else on the site — it is what you are supposed to do when a page moves — so nothing in an ordinary audit flags it. It is only in an hreflang tag that pointing at a redirect is a fault, because the tag is asserting an identity rather than following a link.

A noindexed alternate is a contradiction. The tag says "index this as the German version"; the page says "do not index me". The page wins, and the cluster is short a language.

One broken target can cost the whole cluster. hreflang sets are evaluated as a group. A group with an unreachable member is a group that does not agree with itself.

Fixing it restores the locale targeting the tags were written to provide.

How to fix it

  1. Read the list on the finding. Each entry names the address, the HTTP status it returned, and whether it was rejected for its status or for a noindex directive.

  2. For a redirect, point the tag at the destination. If /de/preise now 301s to /de/tarife, every hreflang naming /de/preise should name /de/tarife instead — on every page of the cluster, not only the one reported.

  3. For a 404 or a 410, decide whether the translation still exists. If it does, point the tag at its current address. If it does not, remove the tag from every page in the cluster; an hreflang set is allowed to have fewer languages than it used to.

  4. For a 5xx, fix the page. A server error is a fault in its own right and this check is only telling you that a second thing depends on it.

  5. For a noindexed page, choose which signal you meant. If the translation is supposed to rank, remove the noindex. If it is deliberately hidden — a staging locale, a thin machine translation — remove the hreflang tags naming it.

  6. Regenerate the hreflang blocks from one source. These tags drift because they are maintained by hand in several templates. Generating all of them from one list of locale-to-URL pairs is what stops it happening again.

Examples

Example 1 — an alternate that redirects

Problematic: /de/preise was renamed and now 301s to /de/tarife, but the hreflang blocks still name the old address.

<!-- https://example.com/en/pricing -->
<link rel="alternate" hreflang="de" href="https://example.com/de/preise" />

Corrected: name the address that answers.

<link rel="alternate" hreflang="de" href="https://example.com/de/tarife" />

Example 2 — an alternate that is gone

Problematic: the Italian translation was retired and /it/chi-siamo now returns 404, but every other language still lists it.

<link rel="alternate" hreflang="it" href="https://example.com/it/chi-siamo" />

Corrected: remove the tag from every page in the cluster. A cluster of two languages is a valid cluster.


Example 3 — an alternate the site has asked not to be indexed

Problematic: the Spanish locale is a machine translation the team is not ready to publish, so it carries a noindex — while the hreflang tags still point at it.

<!-- https://example.com/en/about -->
<link rel="alternate" hreflang="es" href="https://example.com/es/sobre-nosotros" />

<!-- https://example.com/es/sobre-nosotros -->
<meta name="robots" content="noindex" />

Corrected: pick one. Either drop the noindex because the page is meant to rank, or drop the hreflang tags naming it because it is not.

How PixyScan detects this

The status of an hreflang target is a fact about a different page, so this runs once after the crawl rather than while any single page is being read.

  1. Collect every page's hreflang set — the language code and address of each link rel="alternate" the crawler recorded.

  2. Match each declared address to a page in the scan. The comparison is the one the crawl files pages under: fragment dropped, a trailing slash dropped from the path, campaign parameters dropped. A target written with www. or a different scheme falls back to the page the crawl stored, unless two crawled pages share that looser address.

  3. Skip the page's own self-referencing tag.

  4. Read what the crawl found there. A target is reported when either:

    • its HTTP status was recorded and is not 200 — including any 3xx redirect, which is a signpost rather than a page; or
    • it answered 200 and the page carries a noindex directive, in the robots meta tag or the X-Robots-Tag header.
  5. Report the page that declared the tag, listing each unusable address with the status it returned and why it was rejected.

What is deliberately never reported:

  • A target that is not in this scan. A ccTLD sibling on another domain cannot be measured from a crawl scoped to one host.
  • A target whose status was never measured. A page the crawl could not reach — the page budget ran out, the run was cancelled — has no status code, and "we did not look" is not "it is broken".
  • A target with no page analysis recorded. A PDF, a feed, or a page that failed before analysis has nothing to read a noindex directive off.
  • A page that did not itself serve. An error page's own markup is not judged.

When a target is reported here, the missing-return-link check stays silent about the same pair: a page search engines discard has no return link to give, and reporting both would be the same fault counted twice.

What we store

Storage Level

Page Level — the finding is attached to the page whose hreflang set points at the unusable address, because that is where the tag that needs editing lives.


Database Table / Prisma Model

PageInternationalSeo, joined against Url and PageSeoBasicsData.


Fields Used

Field Type Description
PageInternationalSeo.hreflangTags Json The { hreflang, href } pairs declared on the page
Url.url String The address each hreflang target is matched against
Url.statusCode Int? What the crawl received at the target; null means unmeasured
PageSeoBasicsData.indexabilityNoindexAbsent Boolean? False when the target carries a noindex directive

Detection Dependencies

  • HTML Document — the link rel="alternate" hreflang tags in the page head
  • HTTP Response — the status code recorded for each target page
  • The whole scan — the target pages' own rows, which is what makes this a post-crawl join

Further reading