Skip to content
Issue docs

HTML served without text compression

Importanthtml_not_compressedIssue 173

What is this issue?

The page's HTML was sent over the network uncompressed, even though the request said it could accept compressed responses.

Every modern browser (and PixyScan's crawler) sends a header like:

Accept-Encoding: gzip, deflate, br, zstd

A server that honours it compresses the HTML and says so:

Content-Encoding: br

This page's response had no Content-Encoding of gzip, Brotli (br), zstd or deflate, so the full, uncompressed document crossed the network.

For a page to pass this check:

  • A successful (2xx) HTML response carries Content-Encoding: gzip, br, zstd or deflate.

Very small pages are not reported. Below 1 KB, a page fits in the first network round trip either way, and compression saves next to nothing — most servers and CDNs skip it at that size by design.

Why it matters

  • HTML compresses extremely well. Markup is repetitive text; gzip typically shrinks it by 70-80%, and Brotli a little more. An uncompressed 100 KB page is about 20 KB compressed.

  • Every page view pays for it. The HTML is the first thing downloaded and nothing else can start until it arrives. Extra bytes here delay First Contentful Paint and Largest Contentful Paint directly, especially on mobile networks.

  • Core Web Vitals. Slower HTML delivery shows up in Google's page experience signals, and Lighthouse reports it as "Enable text compression".

  • Crawl efficiency. Search engine crawlers accept compressed responses too. Smaller responses let them fetch more of the site in the same time.

  • It is usually one setting. Compression is built into every mainstream web server and CDN; when it is off, it is almost always a configuration oversight.

Effect on the health score

This is an important issue. It deducts from the Delivery & Trust score for each affected page.

How to fix it

Turn on compression for HTML (and, while you are there, CSS, JavaScript, JSON and SVG).

nginx

gzip on;
gzip_types text/html text/css application/javascript application/json image/svg+xml;
gzip_min_length 1024;
## With the Brotli module installed:
brotli on;
brotli_types text/html text/css application/javascript application/json image/svg+xml;

text/html is always compressed once gzip on is set; the list adds the other types.

Apache

AddOutputFilterByType DEFLATE text/html text/css application/javascript application/json
## Apache 2.4.26+ with mod_brotli:
AddOutputFilterByType BROTLI_COMPRESS text/html text/css application/javascript

Node.js (Express)

import compression from 'compression'
app.use(compression())

CDNs: Cloudflare, Fastly, CloudFront and Akamai can compress at the edge. On CloudFront, enable Compress objects automatically on the cache behaviour.

Check the result:

curl -sI -H 'Accept-Encoding: gzip, br' https://example.com/ | grep -i content-encoding

If the origin compresses but the response still arrives uncompressed, a proxy or CDN in front of it is decompressing it — check that layer's settings.

Examples

Example 1: Uncompressed HTML

HTTP/1.1 200 OK
Content-Type: text/html; charset=utf-8
Content-Length: 48211

Reported: 48 KB of HTML with no Content-Encoding.

Example 2: Brotli

HTTP/2 200
content-type: text/html; charset=utf-8
content-encoding: br

Passes.

Example 3: identity

HTTP/1.1 200 OK
Content-Type: text/html
Content-Encoding: identity

Reported: identity means "not encoded".

Example 4: A tiny page

HTTP/1.1 200 OK
Content-Type: text/html
Content-Length: 612

Passes: under 1 KB, compression saves nothing worth having.

Example 5: A PDF

HTTP/1.1 200 OK
Content-Type: application/pdf

Not checked: this check is about HTML documents.

How PixyScan detects this

  1. Asks for compression the way a browser does. The crawler's requests advertise gzip, deflate and Brotli (and zstd on the default engine), so a response without compression is the server's choice.

  2. Reads the encoding the server actually sent. The crawler's HTTP client decompresses responses itself and then removes Content-Encoding from the headers it hands back. PixyScan reads the raw header lines the server sent instead, so a compressed page is never mistaken for an uncompressed one.

  3. Only judges pages where the answer means something:

    • the response headers were actually read,
    • the status is 2xx (error pages are reported elsewhere),
    • the Content-Type is HTML (text/html or application/xhtml+xml),
    • the document is at least 1,024 bytes — measured from the declared Content-Length, or from the HTML itself when there is none.
  4. Accepts any compressing coding: gzip, br, zstd, deflate, compress, and the x-gzip / x-compress aliases, in any case. A list such as identity, gzip counts as compressed.

  5. Reports anything else, including identity (which means "none") and unrecognised values. The finding records the size and the encoding found.

  6. Applies to HTTP and HTTPS pages alike.

What we store

Storage Level

Page Level


Database Table / Prisma Model

PageResponseHeader, and the finding on audit_issues.details


Fields Used

Field Type Description
page_response_headers.headers Json The page's response headers. content-encoding is stored as the server sent it, recovered from the raw header lines when the HTTP client removed it after decompressing

Stored Fields on the finding

Field Type Description
message String The size of the document and the encoding found
htmlBytes Int Size of the HTML: the declared Content-Length, or the document's byte length when none was sent
contentEncoding String The Content-Encoding sent, or null when there was none

Detection Dependencies

  • HTTP Response headers, including the raw header lines
  • HTTP status code
  • Content-Type

Further reading