Skip to content
Issue docs

URL containing non-ASCII characters

Suggestionurl_non_ascii_charactersIssue 119

What is this issue?

This check reports a page whose address contains characters outside the ASCII range -- accented letters, CJK characters, emoji, anything above U+007F.

It examines the path only. The hostname is excluded: an internationalised domain is registered and requested in its punycode form (xn--...), so it is already ASCII by the time anything reaches it, and the owner of such a domain cannot change that.

A character that is percent-encoded is still a non-ASCII character. /caf%C3%A9 and /café are the same address written two ways, and both are reported. Percent-encoding of an ASCII character is not: %20 is a space, and spaces in URLs are covered by a separate check.

The query string and the fragment are excluded too. A non-ASCII value there is usually somebody's search term rather than an address anyone chose to publish, and the pages carrying them are the search and filter pages a site already keeps out of the index.

For a URL to pass:

  • Its path contains only ASCII characters.

Why it matters

Search engines handle non-ASCII URLs correctly, and Google has said so explicitly. The cost is in every other system that touches the address:

  • Inconsistent encoding: browsers display the decoded form, copy the encoded form, and different tools encode differently. One page ends up linked, reported and bookmarked under several spellings.
  • Broken links: email clients, chat apps and older CMSs still truncate or mangle percent-encoded sequences, so the link that was pasted is not the link that arrives.
  • Analytics noise: the encoded and decoded forms often appear as separate rows.
  • Readability: a URL that reads as /caf%C3%A9-guide in a search result or a status bar tells the reader nothing.

Reported as a suggestion, so it never deducts from the health score. Nothing about the page stops working, Google says plainly that it handles these URLs, and the fix is a slug change plus a permanent redirect rather than a cheap edit. It also reaches every page of a site written in a non-Latin script, where the spelling is correct and there is nothing to fix at all.

How to fix it

  1. Transliterate the slug. Generate URL slugs by converting accented and non-Latin characters to their closest ASCII equivalent: café becomes cafe, Über becomes uber.

  2. Fix it in the slug generator, not page by page. Whatever turns a title into a URL is where this originates, and correcting it there stops new pages inheriting the problem.

  3. Redirect the old address. Every URL you change needs a permanent (301) redirect from the encoded form.

  4. Update internal links and the sitemap to the new ASCII form, so nothing keeps advertising the old one.

  5. Keep the language in the content, not the URL. The page's own language is declared with lang and hreflang; it does not need to be spelled out in the path.

  6. Leave an internationalised DOMAIN alone. If your domain itself is an IDN that is a deliberate choice and is not what this check reports.

Examples

Example 1: An ASCII slug

Scenario: A French-language page with a transliterated slug.

Passes because: every character in the path is ASCII.

https://example.com/blog/cafe-guide

Example 2: An accented character in the path

Scenario: A slug generated straight from the page title.

Fails because: the path contains a character above the ASCII range, whether it is written literally or percent-encoded. Both spellings below are the same address.

https://example.com/blog/cafe-guide  (correct)
https://example.com/blog/caf%C3%A9-guide  (reported)

Corrected version: transliterate the slug and 301-redirect the old form.

https://example.com/blog/cafe-guide

Example 3: An encoded space

Scenario: A file name with a space in it.

Passes this check because: %20 decodes to a space, which is ASCII. The space is still worth fixing, and a separate check reports it.

https://example.com/files/annual%20report.pdf

Example 4: A non-ASCII search term

Scenario: An internal search results page.

Passes because: only the path is judged. The value is what somebody typed into a search box, not an address the site chose to publish.

https://example.com/search?q=caf%C3%A9

How PixyScan detects this

  1. Parses the page's final address and takes its PATH. The hostname, the query string and the fragment are all deliberately left out -- the hostname because an internationalised domain is already punycode by the time anything sees it, and the other two because a non-ASCII value there is usually somebody's search term rather than an address the site chose to publish.

  2. Looks for two things. A character written literally above U+007F, and a percent-escape of a byte at or above 0x80 -- %C3, %E5 and so on, which is what a non-ASCII character becomes once a URL has been parsed.

  3. Ignores escapes below 0x80. %20 is a space and %2F is a slash; both are ASCII, and neither is reported here.

  4. Reports a truncated escape too. A sequence such as /%E5%AF cannot be decoded, but %E5 is still a byte above 0x7F, so the URL is reported rather than skipped. A URL that is too broken to decode is not evidence that nothing is wrong with it.

  5. Falls back to the whole string when the address cannot be parsed at all. There is no path to isolate in that case, and a URL too broken to parse is not evidence that nothing is wrong with it.

  6. Names the address on the finding, in the form the crawler fetched.

What we store

Storage Level

Page Level


Database Table / Prisma Model

audit_issues.details


Stored Fields

Field Type Description
url String The page URL that was examined
message String What was found and what is required instead

Detection Dependencies

  • HTTP Response
  • The page URL as the crawler fetched it

Note

The evidence is kept on the finding rather than in a per-page table. The address is already stored on the Urls row, in the percent-encoded form the crawler requested.

Further reading