Skip to content
Issue docs

Thin content: fewer than 200 words of body text

Importantthin_contentIssue 127

What is this issue?

This check counts the words of body text on the page and reports it when the total is below 200 words.

Body text means the page's own prose. The site furniture — the navigation, the header and the footer — is removed before counting, and so are script and style blocks and anything the markup itself declares hidden (a hidden attribute, an inline display: none, a collapsed drawer). A site-wide navigation is the same few hundred words on every page; counting it would give every page the same baseline and make a genuinely empty page look exactly like a full one.

For a page to pass this check:

  • Its own prose runs to 200 words or more.

Example: an article of 900 words passes. A product page whose only text is a name, a price and a "Add to basket" button does not — even if the surrounding template carries a thousand words of menu, promo and footer links.

The 200-word threshold is the same floor PixyScan already applies before it will report a reading grade for a page: below it there is not enough prose for the measurement to mean anything.

Why it matters

A page with almost no text has almost nothing for a search engine to match a query against. That has two concrete consequences:

  • It may never be indexed. "Crawled — currently not indexed" in Search Console is the state Google puts low-value pages in. The page was fetched, read and set aside. It cannot rank for anything from there.
  • It competes badly when it is indexed. Ranking is a comparison. A page with forty words is being compared against pages that answer the question, and it loses that comparison on the merits.

It matters most in bulk. A template that produces thin pages produces hundreds of them — a tag archive per tag, a location page per postcode, a variant page per size — and a site made mostly of pages Google declines to index is a site whose crawl budget is being spent on nothing.

Not every short page is a fault. A contact page, a login screen or a thank-you page is correct at forty words, and this check cannot tell one of those from an unfinished article. That is why a page carrying noindex is not reported at all: noindexing a page that is meant to be short is one of the fixes below, and a page that took the advice must not still be reported for taking it.

What is left after that is the case this check is for — an indexable page with nothing on it — and that is graded IMPORTANT, because it is precisely the definition: the page can be indexed and competes worse than it should. Fixing it lifts the health score in proportion to how much of the site was affected.

How to fix it

  1. Decide whether the page should exist. A thin page that duplicates what a better page already covers is not a writing job, it is a deletion or a redirect. Merging three thin pages into one substantial one is almost always the right move.

  2. Answer the question the page promises. If the URL says "size guide", the page owes a size guide. Add the specifications, the measurements, the caveats and the questions people actually ask — not padding.

  3. Fix it in the template when it is a template. Hundreds of thin pages are never hundreds of writing tasks. Find the template that generates them and give it something real to render: specifications on a product page, a description and a count on a category page.

  4. Noindex the pages that are meant to be short. A login screen, a basket, a thank-you page and a print view are correct at thirty words and have no business in the index. noindex says so, and takes them out of this comparison for good.

  5. Do not pad. Text written to clear a word count reads like text written to clear a word count, and it is exactly what Google's helpful-content guidance is aimed at. A short page that answers the question completely is better than a long one that does not.

Examples

Example 1: A product page with nothing on it

Scenario: An e-commerce template renders a name, a price and a button. The surrounding page carries a large menu and a large footer.

Fails because: the site furniture is not counted, and the page's own prose is eight words.

<body>
  <nav><!-- 400 words of menu --></nav>
  <main>
    <h1>Trail Runner 3</h1>
    <p>£129.00</p>
    <button>Add to basket</button>
  </main>
  <footer><!-- 300 words of links --></footer>
</body>

Corrected version: the template renders what the buyer needs to decide.

<main>
  <h1>Trail Runner 3</h1>
  <p>£129.00</p>
  <p>A 280 g trail shoe with a 6 mm drop… <!-- materials, fit, drainage,
  outsole, sizing advice, care --></p>
</main>

Example 2: A tag archive with one entry

Scenario: A blog generates a page for every tag. Half of the tags have been used once.

Fails because: the page is a heading and a single link.

Corrected version: stop generating a page for a tag used fewer than three times, and noindex the archives that remain thin.

Example 3: A page that is short on purpose

Scenario: A thank-you page shown after a form submission.

Fails the word count, and should not be in the index at all. The fix is <meta name="robots" content="noindex">, not more words.

How PixyScan detects this

  1. Takes the page's own prose. From the HTML the crawler read — the response body on the plain HTTP engine, the document after scripts have run on the JavaScript engine — PixyScan removes the elements that are not this page's content: script, style, noscript and template blocks, anything the markup itself declares hidden, and the nav, header and footer regions. What is left is the text a reader came to the page for.

    Text is read with the element boundaries respected — a heading and the paragraph after it are two pieces of text, not one glued token — while inline markup inside a word (<span>, <em>, a drop cap) is transparent and does not split it.

  2. Counts the words. Hyphens and apostrophes are inside a word, so "state-of-the-art" is one word and "don't" is one word. This is the same count the page report shows as the page's word count, and the same count the reading-grade check works from — one number, computed once.

  3. Compares it against 200. A page with exactly 200 words passes; 199 fails. The finding always names the number it actually applied, because the crawler honours a per-scan thinContentMinWords override when the job carries one. No screen sets it today, so in practice every scan applies 200.

  4. Steps aside in two cases. A short page is not always a defect, and where another finding already covers it, PixyScan does not raise a second row:

    • The page is noindexed. It is out of the index, so its word count cannot cost a ranking — and "noindex the pages that are meant to be short" is one of the fixes recommended below, so reporting a page that took the advice would be incoherent.
    • The page is a soft 404 (#129), and that check is switched on. Every soft 404 is empty by definition, and the fix is to return a 404 status, not to write more copy. If the Indexability lens is off, #129 is never reported and this finding is not stood down for it.
  5. Reports the measurement. The finding carries the words counted and the minimum required, so the shortfall is on the row itself.

The check makes no judgement about what the words say. Whether the prose is any good is a separate question, and the reading-level check is where PixyScan says anything about it.

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageSeoBasicsData — the measurement

audit_issues.details — the finding


Fields Used

Field Type Description
wordCount Int Words of body text on the page, as counted for this check

The finding itself carries the same measurement alongside the threshold that was applied, so a report can be read without joining back to the page row:

Field Type Description
url String The page the text was read from
wordCount Int The words counted
minimumWords Int The minimum this scan applied
actual Int The measured count, structured
expectedMin Int The minimum, structured
difference Int How many words short of the minimum

Detection Dependencies

  • HTML Document (the response body on the plain HTTP engine; the document after scripts have run on the JavaScript engine)

Note

wordCount is stored on the page SEO basics row rather than on the readability row, because thin content is a page-basics question and the readability group can be switched off independently. Both come through the same counter, so the two can never disagree.

Further reading