Low text-to-HTML ratio
What is this issue?
This check divides the bytes of text a reader can see by the bytes of the HTML document, and reports the page when the result is below 10%.
Everything visible counts as text — the navigation and the footer included, because those are text the server really sent and the reader really reads. What does not count is markup: tags, attributes, inline scripts, inline styles, comments and anything the page declares hidden.
For a page to pass this check:
- Visible text is at least 10% of the HTML document's size.
Example: a 40 KB article page with 6 KB of prose is at 15% and passes. A 300 KB page carrying 8 KB of prose and 290 KB of inline application state is at under 3% and does not.
This is advice, not a defect. No search engine reads a text-to-HTML ratio, and a modern framework can put a perfectly good page under 10% on its own. What the number is good for is spotting the page that is all chrome and no content.
A page that is thin (#127) is not reported here as well — a page with almost no text has almost no ratio by arithmetic, and #127 is the finding that names the fix. That holds even when the thin finding itself was suppressed, so a short page never produces two rows for one emptiness. A page reported as a soft 404 (#129) is likewise left to that finding.
Why it matters
It never deducts from the health score, and it should not: there is no ranking penalty for a low ratio, and plenty of excellent pages have one.
What a low ratio is, reliably, is a symptom worth following:
- The content is not in the HTML. The most common cause of a very low ratio on a page that looks full in a browser is that the text arrives after JavaScript runs. On the plain HTTP engine PixyScan reads exactly what the server returned, which is what a crawler sees first — so a low ratio there often means the text a search engine reads is not the text a person reads. (Scanned with JavaScript rendering on, the same page is measured after hydration and the ratio is higher. The two engines answer two different questions about the same URL.)
- The page costs more than it delivers. Every byte is downloaded on a phone, on a train, on a metered connection. A page that is 95% markup is spending a reader's data on the parts of itself nobody came for.
- It surfaces the page whose weight is all framework. A page with a real article on it, buried in twenty times its own size of inline state and dead wrapper markup, has nothing wrong with its word count and everything wrong with what it costs to read. That page is invisible to every other check.
Acting on it does not move the score at all. What it usually surfaces is either a rendering problem worth fixing properly, or a page that has less on it than its owner thinks.
How to fix it
Check the content is in the HTML at all. View the page source, not the rendered DOM. If the article body is not there, the ratio is a symptom and server-side rendering is the fix.
Move inline scripts and styles into files. A stylesheet or a script served from its own URL is cached across the whole site; the same bytes inlined into every page are downloaded again on every page.
Trim the inline state blob. Frameworks serialise the data they hydrate from into the page. Send the fields the page actually renders, not the whole API response.
Remove dead markup. Commented-out sections, abandoned tracking snippets, deeply nested wrapper
divs that exist only to hang a class on, and hidden variants of a component that never render — all of it is downloaded.Then look at whether the page has enough on it. If the markup is already lean and the ratio is still low, the honest reading is that there is not much content on the page. That is the thin-content question, and #127 covers it.
Do not pad the page with text to raise the ratio. The number is a symptom; treating the number rather than the cause makes the page worse.
Examples
Example 1: A real article under a mountain of state
Scenario: A 900-word article renders correctly, and the same page ships 280 KB of serialised API response inline so the framework can hydrate.
Fails because: the prose is 6 KB of a 300 KB document — 2%. The word count is fine, so no other check sees anything.
<body>
<main><h1>Choosing running shoes</h1><p>…900 words…</p></main>
<script>window.__STATE__ = { /* 280 KB of serialised API response */ }</script>
</body>Corrected version: serialise the fields the page actually renders.
<body>
<main><h1>Choosing running shoes</h1><p>…900 words…</p></main>
<script>window.__STATE__ = { /* 4 KB: the fields this page renders */ }</script>
</body>(A page whose article is NOT in the served HTML at all is reported as thin content (#127) instead — see Example 3.)
Example 2: Inline styles repeated on every page
Scenario: A build step inlines the whole stylesheet into every page.
Fails because: 120 KB of CSS is markup, and it is downloaded again on every page rather than cached once.
Corrected version: inline only the critical rules and link the rest.
<style>/* ~2 KB critical CSS */</style>
<link rel="stylesheet" href="/assets/app.css">Example 3: A page that is simply short
Scenario: A lean template with 40 words of prose.
Reported as thin content (#127), not here. The ratio finding steps aside so one page produces one row, and the fix is to give the page something to say.
How PixyScan detects this
Measures the HTML. The size, in bytes, of the document the crawler read. On the plain HTTP engine that is the response body; on the JavaScript engine it is the document after scripts have run, which is the version a rendering crawler sees. The two can differ substantially on a client-rendered page, and the measurement is of whichever one this scan actually read.
Measures the visible text. PixyScan removes the elements a reader never sees —
script,style,noscript,template, and anything the markup itself declares hidden — then takes the text of what is left, collapses runs of whitespace and measures that in bytes.Navigation, header and footer text is kept. It is text the server sent and the reader reads, so it belongs in a ratio of text to markup. (The thin-content check removes it, because there the question is what this page's own prose amounts to. Two questions, two measurements.)
Divides one by the other. Bytes, not characters, on both sides — the question is how much of what was downloaded was content, and a page in a non-Latin script would otherwise measure a third of the ratio it has.
Compares against 10%. Exactly 10% passes; anything below fails.
Steps aside for a thin page. If the page is below the thin-content word count at all, this finding is not raised: a page with almost no text has almost no ratio by arithmetic, so the two would be one fault under two names. That holds even when the thin finding itself was suppressed — a noindexed short page already has its answer, and this row must not become the second one.
What we store
Storage Level
Page Level
Database Table / Prisma Model
audit_issues.details
Fields Used
| Field | Type | Description |
|---|---|---|
| url | String | The page that was measured |
| textBytes | Int | Bytes of visible text |
| htmlBytes | Int | Bytes of the HTML document the crawler read |
| ratio | Float | textBytes / htmlBytes |
| expectedMin | Float | The floor applied, 0.1 |
Detection Dependencies
- HTML Document (the response body on the plain HTTP engine; the document after scripts have run on the JavaScript engine)
Note
There is no per-page column for this measurement. Both numbers are properties of one response rather than facts about the page worth keeping between scans, and a page that passes has nothing worth storing — so the evidence lives on the finding, where the report reads it.