Skip to content
Issue docs

Low text-to-HTML ratio

Suggestionlow_text_to_html_ratioIssue 128

What is this issue?

This check divides the bytes of text a reader can see by the bytes of the HTML document, and reports the page when the result is below 10%.

Everything visible counts as text — the navigation and the footer included, because those are text the server really sent and the reader really reads. What does not count is markup: tags, attributes, inline scripts, inline styles, comments and anything the page declares hidden.

For a page to pass this check:

  • Visible text is at least 10% of the HTML document's size.

Example: a 40 KB article page with 6 KB of prose is at 15% and passes. A 300 KB page carrying 8 KB of prose and 290 KB of inline application state is at under 3% and does not.

This is advice, not a defect. No search engine reads a text-to-HTML ratio, and a modern framework can put a perfectly good page under 10% on its own. What the number is good for is spotting the page that is all chrome and no content.

A page that is thin (#127) is not reported here as well — a page with almost no text has almost no ratio by arithmetic, and #127 is the finding that names the fix. That holds even when the thin finding itself was suppressed, so a short page never produces two rows for one emptiness. A page reported as a soft 404 (#129) is likewise left to that finding.

Why it matters

It never deducts from the health score, and it should not: there is no ranking penalty for a low ratio, and plenty of excellent pages have one.

What a low ratio is, reliably, is a symptom worth following:

  • The content is not in the HTML. The most common cause of a very low ratio on a page that looks full in a browser is that the text arrives after JavaScript runs. On the plain HTTP engine PixyScan reads exactly what the server returned, which is what a crawler sees first — so a low ratio there often means the text a search engine reads is not the text a person reads. (Scanned with JavaScript rendering on, the same page is measured after hydration and the ratio is higher. The two engines answer two different questions about the same URL.)
  • The page costs more than it delivers. Every byte is downloaded on a phone, on a train, on a metered connection. A page that is 95% markup is spending a reader's data on the parts of itself nobody came for.
  • It surfaces the page whose weight is all framework. A page with a real article on it, buried in twenty times its own size of inline state and dead wrapper markup, has nothing wrong with its word count and everything wrong with what it costs to read. That page is invisible to every other check.

Acting on it does not move the score at all. What it usually surfaces is either a rendering problem worth fixing properly, or a page that has less on it than its owner thinks.

How to fix it

  1. Check the content is in the HTML at all. View the page source, not the rendered DOM. If the article body is not there, the ratio is a symptom and server-side rendering is the fix.

  2. Move inline scripts and styles into files. A stylesheet or a script served from its own URL is cached across the whole site; the same bytes inlined into every page are downloaded again on every page.

  3. Trim the inline state blob. Frameworks serialise the data they hydrate from into the page. Send the fields the page actually renders, not the whole API response.

  4. Remove dead markup. Commented-out sections, abandoned tracking snippets, deeply nested wrapper divs that exist only to hang a class on, and hidden variants of a component that never render — all of it is downloaded.

  5. Then look at whether the page has enough on it. If the markup is already lean and the ratio is still low, the honest reading is that there is not much content on the page. That is the thin-content question, and #127 covers it.

Do not pad the page with text to raise the ratio. The number is a symptom; treating the number rather than the cause makes the page worse.

Examples

Example 1: A real article under a mountain of state

Scenario: A 900-word article renders correctly, and the same page ships 280 KB of serialised API response inline so the framework can hydrate.

Fails because: the prose is 6 KB of a 300 KB document — 2%. The word count is fine, so no other check sees anything.

<body>
  <main><h1>Choosing running shoes</h1><p>…900 words…</p></main>
  <script>window.__STATE__ = { /* 280 KB of serialised API response */ }</script>
</body>

Corrected version: serialise the fields the page actually renders.

<body>
  <main><h1>Choosing running shoes</h1><p>…900 words…</p></main>
  <script>window.__STATE__ = { /* 4 KB: the fields this page renders */ }</script>
</body>

(A page whose article is NOT in the served HTML at all is reported as thin content (#127) instead — see Example 3.)

Example 2: Inline styles repeated on every page

Scenario: A build step inlines the whole stylesheet into every page.

Fails because: 120 KB of CSS is markup, and it is downloaded again on every page rather than cached once.

Corrected version: inline only the critical rules and link the rest.

<style>/* ~2 KB critical CSS */</style>
<link rel="stylesheet" href="/assets/app.css">

Example 3: A page that is simply short

Scenario: A lean template with 40 words of prose.

Reported as thin content (#127), not here. The ratio finding steps aside so one page produces one row, and the fix is to give the page something to say.

How PixyScan detects this

  1. Measures the HTML. The size, in bytes, of the document the crawler read. On the plain HTTP engine that is the response body; on the JavaScript engine it is the document after scripts have run, which is the version a rendering crawler sees. The two can differ substantially on a client-rendered page, and the measurement is of whichever one this scan actually read.

  2. Measures the visible text. PixyScan removes the elements a reader never sees — script, style, noscript, template, and anything the markup itself declares hidden — then takes the text of what is left, collapses runs of whitespace and measures that in bytes.

    Navigation, header and footer text is kept. It is text the server sent and the reader reads, so it belongs in a ratio of text to markup. (The thin-content check removes it, because there the question is what this page's own prose amounts to. Two questions, two measurements.)

  3. Divides one by the other. Bytes, not characters, on both sides — the question is how much of what was downloaded was content, and a page in a non-Latin script would otherwise measure a third of the ratio it has.

  4. Compares against 10%. Exactly 10% passes; anything below fails.

  5. Steps aside for a thin page. If the page is below the thin-content word count at all, this finding is not raised: a page with almost no text has almost no ratio by arithmetic, so the two would be one fault under two names. That holds even when the thin finding itself was suppressed — a noindexed short page already has its answer, and this row must not become the second one.

What we store

Storage Level

Page Level


Database Table / Prisma Model

audit_issues.details


Fields Used

Field Type Description
url String The page that was measured
textBytes Int Bytes of visible text
htmlBytes Int Bytes of the HTML document the crawler read
ratio Float textBytes / htmlBytes
expectedMin Float The floor applied, 0.1

Detection Dependencies

  • HTML Document (the response body on the plain HTTP engine; the document after scripts have run on the JavaScript engine)

Note

There is no per-page column for this measurement. Both numbers are properties of one response rather than facts about the page worth keeping between scans, and a page that passes has nothing worth storing — so the evidence lives on the finding, where the report reads it.

Further reading