Character Encoding - ASCII, ISO-8859-1, UTF-8, UTF-16

ASCII, ISO-8859-1, UTF-8, UTF-16

Two ideas hide behind the phrase character encoding, and most encoding trouble comes from mixing them up. A character set gives every character a number, called a code point: Unicode says the euro sign is U+20AC. An encoding says how that number is written as bytes: UTF-8 writes U+20AC as the three bytes E2 82 AC. Unicode is the set. UTF-8, UTF-16 and UTF-32 are encodings of it.

ASCII is the oldest of these still in use. It is a 7-bit set of 128 code points: 95 printable characters — the unaccented English letters, the digits and common punctuation — and 33 control codes such as tab and newline. No accented letters, and no currency symbol but the dollar.

ISO-8859-1 (Latin-1), Windows-1252 and the rest of the ISO/IEC 8859 family use the eighth bit to add 128 more code points. All of them keep ASCII unchanged in the lower half, and all of them differ from one another in the upper half: ISO-8859-1 fills it with Western European letters, ISO-8859-5 with Cyrillic, ISO-8859-7 with Greek. Windows-1252 is ISO-8859-1 with printable characters at 0x80–0x9F, where ISO-8859-1 has control codes — that is where the curly quotes, the em dash and the euro sign sit. The two are mislabelled for each other so often that HTML now requires a browser to read a page marked ISO-8859-1 as Windows-1252.

Unicode replaces all of them with one set covering every writing system. It is not a 16-bit scheme — that was true of Unicode version 1 and the belief has outlived it by thirty years. Since 1996 the code space has run from U+0000 to U+10FFFF: 1,114,112 code points, more than 150,000 of them assigned so far. The first 128 are ASCII's and the first 256 are ISO-8859-1's, which is all that "Unicode is backwards compatible" means — the numbers match, not the bytes.

Three encodings turn those numbers into bytes.

UTF-8 is variable width, one to four bytes per character. ASCII characters keep their single byte and their value, so a file of plain English text is byte-for-byte identical in ASCII and in UTF-8. Six bytes is a figure still widely quoted; it came from the original ISO 10646 definition, which reached far past U+10FFFF, and RFC 3629 capped UTF-8 at four bytes in 2003.

UTF-16 uses two bytes for most characters and four for the rest, through the surrogate pairs below. UTF-32 uses four bytes for every character without exception, which makes it simple to index and wasteful to store, so it is rare outside a program's memory. On the web the answer is UTF-8, and has been for years.

Surrogate Pairs

UTF-16 has 16 bits in a code unit and Unicode outgrew them. A code point above U+FFFF — emoji, the rarer CJK characters, most historic scripts — is written as a surrogate pair: a high surrogate in the range D800–DBFF followed by a low surrogate in the range DC00–DFFF. Both ranges are permanently reserved and stand for nothing on their own, so a pair can never be confused with a real character.

Surrogates belong to UTF-16 alone. UTF-8 and UTF-32 have no use for them and never contain one.

Converting between a code point and its two halves, in JavaScript:
S = 0x10000 + (H - 0xD800) * 0x400 + (L - 0xDC00);
H = Math.floor((S - 0x10000) / 0x400) + 0xD800;
L = ((S - 0x10000) % 0x400) + 0xDC00;

Character Encoding in HTML

Use UTF-8, and declare it. A browser takes the first answer it finds, in this order: the charset in the HTTP Content-Type header, a byte order mark, the meta declaration in the document, and then a guess. The guess is a legacy encoding chosen by the browser's locale — Windows-1252 across most of the West — and it is where mojibake begins.

<meta charset="utf-8"> is the HTML5 form. <meta http-equiv="Content-Type" content="text/html; charset=utf-8"> is the older one and still works. Either has to sit within the first 1024 bytes of the file, so put it immediately after <head>: the browser has to start parsing to find the declaration, and whatever it read first may have to be thrown away and read again.

Do not serve HTML as UTF-16. A meta element cannot announce it — the parser must be able to read that tag as ASCII to get that far — so it depends on a byte order mark or on the HTTP header, and it buys nothing over UTF-8.

Mojibake is the classic symptom: a page shows ’ where it should show a right single quotation mark ’ (U+2019). The file is usually fine. It is UTF-8, in which that character is the three bytes E2 80 99, and it is being read as Windows-1252, which shows each of those bytes as a character of its own. Every non-ASCII character does something similar, and two or three Latin letters and symbols standing where one character belongs is the signature. The fix is to declare the encoding the file actually uses — not to replace the character with a plain apostrophe, which throws away the typography to work around a label.

Byte Order Mark

A byte order mark is the code point U+FEFF written at the very start of a file. In UTF-16 it says which of the two byte orders is in use, which is what it is for. UTF-8 has no byte order to mark, so there the mark is neither required nor recommended — but some editors add it anyway, as the three bytes EF BB BF.

Those three bytes are why a BOM breaks PHP. They are content, sent before anything the script prints, so the first call to header() or session_start() fails with "headers already sent". Save PHP files as UTF-8 without a BOM; any editor worth using offers the choice.

Character Entities

In HTML, &amp; for & and &lt; for < are the two that always matter, because an ampersand starts an entity and a less-than sign starts a tag. &gt; for > is convention rather than necessity. Inside an attribute value, escape whichever quote delimits it: &quot; in a double-quoted value, &#39; in a single-quoted one.

&apos; is the exception worth remembering. It is defined in XML and in HTML5 and not in HTML 4, so &#39; is the safer way to write an apostrophe. Numeric references — &#8217;, or &#x2019; in hexadecimal — can write any character at all, which is what to reach for when a file's own encoding cannot hold one.

Character Encoding in URLs

Percent-encoding encodes bytes, not characters: a percent sign and two hexadecimal digits, one such triplet per byte. That distinction is the whole of it. The letter w is one byte, 0x77, and encodes as %77, which is why https://en.wikipedia.org/wiki/Main_Page and https://en.wikipedia.org/%77%69%6b%69/%4d%61%69%6e%5f%50%61%67%65 are the same address. A character outside ASCII takes as many triplets as it has bytes: é is two in UTF-8, %C3%A9, and one emoji is four.

So encode the character as UTF-8 first, then percent-encode those bytes. The older standards left the choice of encoding open and recommended UTF-8; the URL standard browsers follow today simply requires it.

A query string is the part of a URL after the question mark, carrying key and value pairs: ?key1=value1&key2=value2. The ?, = and & that separate them are structure and stay as they are. Inside a key or a value the same characters are data and have to be encoded — a value containing a bare ampersand splits the query in two.

A domain name is the exception: it is never percent-encoded. A name in another script, such as президент.рф, is converted instead by IDNA into ASCII labels beginning xn-- — here xn--d1abbgf6aiiy.xn--p1ai — before it reaches DNS at all. Browsers show the readable form in the address bar when the name's script is unambiguous, and the xn-- form when mixing scripts could disguise one name as another.

Character Encoding in XML

An XML declaration such as <?xml version="1.0" encoding="utf-8"?> goes at the very start of the document. It is optional for UTF-8, which is the default, and for UTF-16 with a byte order mark; for anything else it is required. Declaring UTF-8 anyway costs one line and settles the question.

A byte the declared encoding cannot explain is fatal rather than cosmetic — an XML parser is required to stop, which is the "invalid character" error you meet on opening the file. XML predefines five entities, &lt;, &gt;, &amp;, &apos; and &quot;, and no others; anything else must be a numeric reference or be declared in the document type.

Character Encoding in JavaScript

An external script's encoding comes from the charset in the HTTP header that serves it, and failing that the browser assumes the encoding of the document that loaded it. The charset attribute on <script> is obsolete in HTML5 and should not be used. Serve scripts as UTF-8 like everything else and the two cannot disagree.

A JavaScript string is a sequence of UTF-16 code units, and it shows. "\u00e4" is an escape for one code unit and is the same string as "ä", while "\u{1f600}" is the newer escape for a whole code point. length, charCodeAt and fromCharCode all count code units, so every character above U+FFFF counts as two and a string holding one emoji has a length of 2. Use codePointAt, String.fromCodePoint, and for...of or the spread operator — which iterate code points — wherever that matters.

Code points are not the last word either. A flag, a family emoji or an Indic conjunct is several code points that display as one character, and splitting between them cuts it in half. That is a question of grapheme clusters rather than of encoding, and it is why deleting one emoji can take four keystrokes in a text box that counts code units.

Character Encoding in MySQL

Use utf8mb4, not utf8. MySQL's utf8 is an alias for utf8mb3, which stores at most three bytes per character and therefore cannot store anything above U+FFFF at all — emoji and the rarer CJK characters are rejected or truncated on the way in. utf8mb4 is UTF-8 as everyone else means it. Set it on the database, the table and the column, with a matching collation such as utf8mb4_0900_ai_ci.

The connection carries its own encoding and it has to agree, or the bytes are converted twice on the way in and once on the way out. SET NAMES utf8mb4; does it, but setting the charset where the connection is made is better: charset=utf8mb4 in a PDO DSN, or mysqli::set_charset(). Code still calling mysql_query("SET NAMES utf8") has two problems rather than one — the mysql_* functions were removed in PHP 7.

Character Encoding in SQL Server

Use nchar and nvarchar rather than char and varchar: the n types store UTF-16 and hold any character. ntext does too, but it is deprecated and nvarchar(max) replaces it. Prefix string literals with N, or the text is converted to the database's non-Unicode code page before it is ever stored or compared: SELECT column FROM table WHERE column = N'data';

Since SQL Server 2019 char and varchar can hold UTF-8 as well, given a collation whose name ends in _UTF8. That is worth having for mostly-ASCII text, where it can halve the storage a UTF-16 column needs.