URL containing non-ASCII characters
What is this issue?
This check reports a page whose address contains characters outside the ASCII
range -- accented letters, CJK characters, emoji, anything above U+007F.
It examines the path only. The hostname is excluded:
an internationalised domain is registered and requested in its punycode form
(xn--...), so it is already ASCII by the time anything reaches it, and the
owner of such a domain cannot change that.
A character that is percent-encoded is still a non-ASCII character. /caf%C3%A9
and /café are the same address written two ways, and both are reported.
Percent-encoding of an ASCII character is not: %20 is a space, and spaces in
URLs are covered by a separate check.
The query string and the fragment are excluded too. A non-ASCII value there is usually somebody's search term rather than an address anyone chose to publish, and the pages carrying them are the search and filter pages a site already keeps out of the index.
For a URL to pass:
- Its path contains only ASCII characters.
Why it matters
Search engines handle non-ASCII URLs correctly, and Google has said so explicitly. The cost is in every other system that touches the address:
- Inconsistent encoding: browsers display the decoded form, copy the encoded form, and different tools encode differently. One page ends up linked, reported and bookmarked under several spellings.
- Broken links: email clients, chat apps and older CMSs still truncate or mangle percent-encoded sequences, so the link that was pasted is not the link that arrives.
- Analytics noise: the encoded and decoded forms often appear as separate rows.
- Readability: a URL that reads as
/caf%C3%A9-guidein a search result or a status bar tells the reader nothing.
Reported as a suggestion, so it never deducts from the health score. Nothing about the page stops working, Google says plainly that it handles these URLs, and the fix is a slug change plus a permanent redirect rather than a cheap edit. It also reaches every page of a site written in a non-Latin script, where the spelling is correct and there is nothing to fix at all.
How to fix it
Transliterate the slug. Generate URL slugs by converting accented and non-Latin characters to their closest ASCII equivalent:
cafébecomescafe,Überbecomesuber.Fix it in the slug generator, not page by page. Whatever turns a title into a URL is where this originates, and correcting it there stops new pages inheriting the problem.
Redirect the old address. Every URL you change needs a permanent (301) redirect from the encoded form.
Update internal links and the sitemap to the new ASCII form, so nothing keeps advertising the old one.
Keep the language in the content, not the URL. The page's own language is declared with
langandhreflang; it does not need to be spelled out in the path.Leave an internationalised DOMAIN alone. If your domain itself is an IDN that is a deliberate choice and is not what this check reports.
Examples
Example 1: An ASCII slug
Scenario: A French-language page with a transliterated slug.
Passes because: every character in the path is ASCII.
https://example.com/blog/cafe-guideExample 2: An accented character in the path
Scenario: A slug generated straight from the page title.
Fails because: the path contains a character above the ASCII range, whether it is written literally or percent-encoded. Both spellings below are the same address.
https://example.com/blog/cafe-guide (correct)
https://example.com/blog/caf%C3%A9-guide (reported)Corrected version: transliterate the slug and 301-redirect the old form.
https://example.com/blog/cafe-guideExample 3: An encoded space
Scenario: A file name with a space in it.
Passes this check because: %20 decodes to a space, which is ASCII. The
space is still worth fixing, and a separate check reports it.
https://example.com/files/annual%20report.pdfExample 4: A non-ASCII search term
Scenario: An internal search results page.
Passes because: only the path is judged. The value is what somebody typed into a search box, not an address the site chose to publish.
https://example.com/search?q=caf%C3%A9How PixyScan detects this
Parses the page's final address and takes its PATH. The hostname, the query string and the fragment are all deliberately left out -- the hostname because an internationalised domain is already punycode by the time anything sees it, and the other two because a non-ASCII value there is usually somebody's search term rather than an address the site chose to publish.
Looks for two things. A character written literally above
U+007F, and a percent-escape of a byte at or above0x80--%C3,%E5and so on, which is what a non-ASCII character becomes once a URL has been parsed.Ignores escapes below
0x80.%20is a space and%2Fis a slash; both are ASCII, and neither is reported here.Reports a truncated escape too. A sequence such as
/%E5%AFcannot be decoded, but%E5is still a byte above0x7F, so the URL is reported rather than skipped. A URL that is too broken to decode is not evidence that nothing is wrong with it.Falls back to the whole string when the address cannot be parsed at all. There is no path to isolate in that case, and a URL too broken to parse is not evidence that nothing is wrong with it.
Names the address on the finding, in the form the crawler fetched.
What we store
Storage Level
Page Level
Database Table / Prisma Model
audit_issues.details
Stored Fields
| Field | Type | Description |
|---|---|---|
| url | String | The page URL that was examined |
| message | String | What was found and what is required instead |
Detection Dependencies
- HTTP Response
- The page URL as the crawler fetched it
Note
The evidence is kept on the finding rather than in a per-page table. The
address is already stored on the Urls row, in the percent-encoded form the
crawler requested.