zgba 站群
UTF-8000: Unlimited UTF-8

UTF-8000: Unlimited UTF-8

There is nothing special-case-y about the example 22-byte code unit here. It is just a good prototypical example, demonstrating the power of UTF-8000 with multiple start bytes.

There are only two special cases, both of which are inherited from UTF-8: ASCII as is, and 2-byte UTF-8 having 4 mandatory content bits to check against overlong encoding as opposed to 5 for all longer length code units.

Here is anatomical diagram of the example 22-byte code unit from the tldr.

See the glossary for more information on the definitions of the terms.

Byte number four is exciting! It is a continuation byte, a start byte, the final start byte, has content bits, and has only some of the mandatory content bits, which are straddled across the final start byte and first non-start byte.

The main contribution of UTF-8000’s specification is clarity on splitting the highest bits of the first byte of UTF-8 code units into self-synchronization bits and start bits, and then making it clear how to stripe the start bits across the continuation bytes if needed, to achieve arbitrarily large code units.

These terms are ordered somewhat by chronology of first requirement, rather than alphabetically, for convenience.

Terms used within definitions are underlined clickable hyperlinks.

A non-negative integer, aka an unsigned integer.

A sequence of UTF-8000 bytes that encode a single codepoint.

The first, one and only, byte that begins a UTF-8000 code unit.

The self-synchronization prefix of a first byte is either 0 for ASCII or 11 for multi-byte code units.

This term is not synonymous with start byte. A first byte is necessarily a start byte, but not the other way around. It is for this reason that first byte is sometimes also known as first start byte.

Fun observation: because of the self-synchronization prefix 0 the upper hex nibble of ASCII bytes can only be one of 0, 1, 2, 3, 4, 5, 6, 7.

This term is mutually exclusive with continuation byte due to self-synchronization.

A byte beyond the first byte of a multi-byte UTF-8000 code unit.

The self-synchronization prefix of a continuation byte is 10, which is also known as the continuation prefix bits.

Fun observation: because of the self-synchronization prefix 10 the upper hex nibble of continuation bytes can only be one of 8, 9, A, B.

This term is mutually exclusive with first byte due to self-synchronization.

Self-Synchronization Prefix

The highest bits of every UTF-8000 byte that indicate whether it is a first byte or a continuation byte.

The possible self-synchronization prefixes form a prefix-free tree:

This piece of the clever architecture of UTF-8, which UTF-8000 inherits, provides the property of self-synchronization at a byte level: we can instantaneously tell what kind of byte we are looking at, and where it should belong in a code unit, just by looking at these highest bits.

This is most useful when decoding part of a file encoded in UTF-8000. If we randomly seek through the file to an arbitrary byte, we can unambiguously tell whether we are at a first byte whence we can begin decoding a new code unit immediately, or that we are at a continuation byte whence we need to seek a little further on in order to find the next first byte in order to begin decoding. Nor do we have to process any bytes prior to our seek position in order to discover some global state or the context of the byte we have seek-ed to; a first byte is always unambiguously a first byte wherever it appears, which we can deduce by its self-synchronization prefix being either 0 or 11.

This is useful not only for random access, but also for error recovery. Suppose that we are decoding an error-prone stream of UTF-8000 bytes and that whenever when we encounter an error (e.g. a rogue 0xC0 byte) we wish to keep calm and carry on instead of immediately exiting. We can yield Unicode replacement characters U+FFFD � and then await the next first byte, discarding anything in the interim.

See the Wikipedia article for self-synchronizing code for more general info.

These bits are highlighted in bright cyan.

A byte containing one or more start bits. The start bytes exist contiguously at the beginning of a UTF-8000 code unit. The power of UTF-8000 is that we can have multiple start bytes, to achieve arbitrary code unit lengths, to encode arbitrarily large codepoints.

Sometimes it is sensible to colloquially also include ASCII as a start byte when we are talking about the bytes towards the start of a code unit, even though ASCII bytes have no start bits.

Every non-ASCII code unit has at least one start byte. The first start byte is the first byte, and it is followed by zero or more continuation bytes that are also start bytes. Therefore because a UTF-8000 code unit can have multiple start bytes, this term is not synonymous with first byte.

In restricting to only UTF-8 without UTF-8000, this term is synonymous with first byte. This is because UTF-8-length code units only require one start byte, whether using up to 4 bytes in the current UTF-8 standard (RFC 3629 (2003)), or using up to 6 bytes in former standards (RFC 2044 (1996) and RFC 2279 (1998)).

The unary-code sequence of bits contained in the start bytes of a multi-byte UTF-8000 code unit that tells us the length of the code unit in bytes.

For a code unit made of n bytes the start bits are n-2 1 bits followed by a terminating 0 bit. To be clear, the start bits include this terminating zero bit. Thus the start bits sequence is of length n-1 and looks like 111…10.

The possible start bits sequences form a prefix-free tree:

For an n byte code unit where n < 8 the start bits all fit together snugly in the first byte. Otherwise they are striped across as many of the first few bytes as they need, filling the free bits that are not occupied by continuation prefix bits.

This is another piece of the clever architecture of UTF-8, which UTF-8000 inherits, that provides the property of self-punctuation also known as a prefix code or a prefix-free code: when decoding a multi-byte code unit, once we have read to the end of the start bytes, that is we have encountered the terminating 0 bit, we know exactly how many bytes we expect in that code unit. Notwithstanding errors we can therefore succeed in decoding the code unit by reading exactly that many bytes, and no more.

This avoids a problem of dumber variable-length encodings whose code units do not intrinsically indicate their length: one has to read beyond the last byte of a code unit, that is one reads the first byte of the next code unit, in order to know that the current code unit has finished. For very dumb encodings which have neither self-synchronization nor self-punctuation, to make random access possible one would have to put dedicated auxiliary bytes, punctuation like a comma byte, between code units to be able to tell where one ends and another begins.

See the Wikipedia articles for prefix code and unary coding for more general info.

This term is mutually exclusive with content bits.

These bits are highlighted in bright magenta.

A byte containing one or more content bits.

A byte being a content byte does not imply that it is a continuation byte. For example a 3-byte code unit begins with 1110xxxx, which contains 4 content bits and is not a continuation byte.

A byte being a continuation byte does not imply that it is a content byte. For example a 22-byte code unit contains 10111111 as its second byte, which is a continuation byte and has no content bits.

The sequence of bits in a code unit beyond the start bits and to the end of the code unit, in which the codepoint’s binary bits are stored. For example a 3-byte code unit, which has the form 1110xxxx 10xxxxxx 10xxxxxx, has 16 content bits.

For ASCII there are 7 content bits. These seven bits xxxxxxx combined with a byte’s highest bit being set to the self-synchronization prefix 0 means that ASCII is perfectly included into UTF-8 without being altered. Thus ASCII code units take the form 0xxxxxxx.

Otherwise for an n byte code unit, where n > 1, there are 5n+1 content bits. This is how we arrive at that formula: We start with n blank bytes, each of which has 8 bits. For each byte 2 bits are taken by the self-synchronization prefix. Then an additional n-1 bits are taken by the start bits. Thus there are 8n - 2n - (n-1) = 5n+1 bits left for content bits. Another way to think about the 5 in this formula is by extending from n-1 bytes to n bytes by appending another continuation byte. By doing this we gain 6 free bits in the continuation byte, but we lose 1 bit to the longer start bits sequence, thus overall we gain 6-1 = 5 bits for content bits.

This term is mutually exclusive with start bits.

These bits are highlighted in lime.

Mandatory Content Byte

A byte containing one or more mandatory content bits.

These are the bytes we check for overlong encoding when decoding a code unit.

Mandatory Content Bits

The first 0, 4, or 5 content bits of a code unit in which there must be at least one 1 bit, lest the bytes form an overlong encoding, which is forbidden.

For ASCII there are 0 mandatory content bits, and thus no anti-overlong checking is required. This is because ASCII is the smallest possible code unit.

For 2-byte UTF-8000 there are 4 mandatory content bits. This is because in the jump from 1-byte ASCII to 2-byte UTF-8 we jump from 7 content bits to 11 content bits. Thus the number of content bits we gain is 11 minus 7 which is 4.

Otherwise for n byte UTF-8000, where n > 2, there are 5 mandatory content bits. This is because in the jump from n-1 byte UTF-8000 to n byte UTF-8000 we add on an extra continuation byte, which has 6 free bits, but we lose 1 bit to the longer start bits sequence. Thus overall the number of content bits we gain is 6 minus 1 which is 5.

Read about overlong encoding for why mandatory content bits are of interest.

These bits are highlighted in bright lime.

Forbidden encodings of codepoints that could be encoded correctly in UTF-8000 using a shorter code unit.

For example one could incorrectly try to encode the codepoint 0x41, 65, ASCII capital A, using 2-byte UTF-8 as 11000001 10000001. Observe that all the mandatory content bits are 0 which is the definition an overlong encoding. This indicates that we could have encoded 0x41 in a shorter code unit, in this case as ASCII 01000001.

Security is one main reason why we forbid overlong encoding. For example we ensure that 11100000 10000000 10000000 cannot be decoded as codepoint 0, the null byte, lest one speciously pass such an overlong byte (code unit) to C functions like strcpy(3) and friends. strcpy would not interpret this code unit as a null byte, leading to a segfault at best, and serious vulnerabilities at least-worst.

Uniqueness of encoding is another reason why we forbid overlong encoding. Every codepoint has one unique valid representation as a UTF-8000 code unit, which is easy to encode and decode using bitshifting.

Fun observation: because all 4 of 2-byte UTF-8’s mandatory content bits lie in the first-and-final start byte, we can explicitly rule out 11000000 (0xC0) and 11000001 (0xC1) as permanently invalid bytes. They will never ever appear anywhere in a valid UTF-8000 code unit!

Many of these properties of UTF-8000 are explained in detail in an appropriate section of the glossary and hyperlinks to the glossary are provided.

The number of content bits and mandatory content bits are very predictable as a function of n, the length of a code unit.

As stated in the tldr, there are only two special cases, both of which are inherited from UTF-8:

1-byte UTF-8 (ASCII) which has two points of interest:

2-byte UTF-8 which has one point of interest:

The remarkable fact that UTF-8000 does not introduce any new special cases in extending UTF-8 is confirmation to me that this is the canonical, correct way to extend UTF-8. In other words UTF-8 in its current restricted 4 byte form is UTF-8000, but only a small part of it.

The fact that we are even able to extend in the first place is also testament to the clever planning and care that Ken Thompson and Rob Pike put into the architecture of UTF-8, which we ensure to maintain as we extend to UTF-8000. Unary code co

View original article