Character encoding declared as something other than UTF-8
What is this issue?
This issue fires when a page does declare a character encoding, but it is something other than UTF-8 — ISO-8859-1, windows-1252, ASCII, Shift_JIS and so on.
<!-- Raises this issue -->
<meta charset="ISO-8859-1" />
<!-- Passes -->
<meta charset="utf-8" />This is the "the declared value is wrong" half of what used to be a single check titled Character encoding missing or not UTF-8. The finding records the declared value under charsetValue, so you can see what the page actually claims.
Why it matters
A legacy encoding cannot represent most of Unicode. Every character outside its range is lost or mangled on the way to the browser — not merely displayed oddly, but unrecoverable from the rendered page.
- It limits what you can publish. A page declared
ISO-8859-1cannot carry a curly quote, an em dash, an emoji, or any non-Latin script. - It corrupts search snippets. Titles and meta descriptions are read from the page using the declared encoding, so mojibake propagates into search results.
- UTF-8 is the web's default. The HTML standard requires authors to use it, and every browser, database and HTTP client handles it natively.
How to fix it
Change the declaration:
<meta charset="utf-8" />Change the declaration and the bytes together. This is the step that is easy to miss: if the file really is encoded as windows-1252, relabelling it as UTF-8 turns correct output into mojibake. Convert the source files (
iconv -f windows-1252 -t utf-8), then relabel.Check the layers underneath, since any one of them can reintroduce the old encoding:
- the HTTP
Content-Typeheader, which overrides the meta tag - your database and its connection charset (
utf8mb4on MySQL, notutf8) - your template engine's and framework's output encoding
- the HTTP
Verify. Re-scan, then load a page with non-ASCII content and confirm it renders correctly.
Examples
1. A legacy Latin-1 declaration
Problem — ISO-8859-1 cannot represent most non-Western characters, so anything outside its range renders as mojibake:
<meta charset="ISO-8859-1" />Café → CaféFixed — re-save the file as UTF-8 and declare it:
<meta charset="utf-8" />Changing only the declaration is not enough. The bytes on disk have to be UTF-8 too, which is what makes this a different job from #28's missing line.
2. Wrong encoding via the http-equiv form
Problem — the declaration is in the older form, and the value is windows-1252:
<meta http-equiv="content-type" content="text/html; charset=windows-1252" />Fixed:
<meta charset="utf-8" />3. Casing and whitespace do not matter
All of these pass. The declared value is trimmed and upper-cased before comparison, so only a genuinely different encoding fails.
Passes:
<meta charset="utf-8" />
<meta charset="UTF-8" />
<meta charset=" UTF-8 " />Fails:
<meta charset="utf8" />utf8 without the hyphen is not the registered name, and it is recorded under charsetValue so the report can say which spelling the page used.
How PixyScan detects this
Value extraction. The crawler reads the declared encoding from either
<meta charset>or<meta http-equiv="content-type">. If neither declares anything, this check does not apply — that ischarset_declaration_missing.Comparison. The declared value is trimmed, upper-cased and compared to
UTF-8, soutf-8,UTF-8and" UTF-8 "all pass.Pass/fail.
- Fails for any other value. The finding records it under
charsetValueand names it in the message. - Passes for UTF-8 in any casing.
- Fails for any other value. The finding records it under
What we store
Storage Level
Page Level — evaluated for each crawled URL.
Database Table / Prisma Model
PageHtmlHeadAudit
One row per crawled URL, keyed by urlId.
Fields Used
| Field | Type | Description |
|---|---|---|
| charsetValue | String? | The encoding the page declares, stored as written — ISO-8859-1, windows-1252, utf-8. This is the value the check compares |
Stored verbatim rather than normalised, so the report can name what the page actually claims instead of only saying that it is wrong.
A null here means nothing was declared, which is charset_declaration_missing (#28) and not this issue. The two conditions cannot both be true of one page.
Detection Dependencies
- HTML Document — the declared encoding is read from
<meta charset>or from thecharset=parameter of<meta http-equiv="content-type">