Draft:UCE-8
| Alias(es) | Unicode Compact Encoding |
|---|---|
| Standard | None (experimental) |
| Current status | Experimental; not adopted by any standards body |
| Classification | Unicode Transformation Format-like, variable-width |
| Extends | ASCII |
| Transforms / Encodes | Unicode |
UCE-8 (Unicode Compact Encoding) is an experimental variable-width character encoding for Unicode that represents each code point in one, two or three bytes. It was designed to reduce the storage and transmission cost of scripts that UTF-8 encodes in three bytes, principally the writing systems of South Asia, Southeast Asia, East Asia and the Horn of Africa.
Unlike UTF-8, in which the leading byte identifies both the length and the general range of a character, UCE-8 places the identifying information in the final byte of a sequence. The high bit of each byte acts solely as a continuation flag, and the terminating byte doubles as an index into one of the encoding's character "pages". Because the two encodings accept almost disjoint sets of byte sequences, a UCE-8 stream and a UTF-8 stream can be held in the same directory or data stream without a label or byte-order mark, a reader distinguishing them by attempting each decoding in turn.
Design
[edit]UCE-8 assigns byte lengths by script rather than by position in the Unicode code space. ASCII and the European and Middle Eastern scripts are unchanged; the two categories below them each shorten by one byte:
| Category | UTF-8 | UCE-8 |
|---|---|---|
| ASCII text | 1 byte | 1 byte |
| European and Middle Eastern scripts | 2 bytes | 2 bytes |
| Asian and African scripts | 3 bytes | 2 bytes |
| Historic scripts and other supplementary-plane characters | 4 bytes | 3 bytes |
The saving applies only to the characters the two-byte tier actually holds; the scripts concerned, and how completely each is covered, are listed under Page map below.
Byte structure
[edit]The most significant bit of every byte indicates whether another byte follows. A byte with the high bit set ("lead byte") is followed by at least one more byte; a byte with the high bit clear ("trail byte") terminates the character. The encoding is therefore self-terminating and requires no length field.
| Tier | Byte pattern | Length | Characters reachable | Range covered |
|---|---|---|---|---|
| 1 | 0xxxxxxx | 1 byte | 128 | ASCII |
| 2 | 1xxxxxxx 0xxxxxxx | 2 bytes | 8,704 | 68 pages of 128 characters |
| 3 | 1xxxxxxx 1xxxxxxx 0xxxxxxx | 3 bytes | 1,114,112 | All remaining code points |
The capacity of the three-byte tier, 128 × 128 × 68, equals 1,114,112, which is also the total size of the Unicode code space (17 × 65,536).
Restricted trail bytes
[edit]Because the trail byte always has its high bit clear, its value falls in the 0–127 range and is therefore always an ASCII byte. Of the 128 possible trail-byte values, 60 are excluded: the 33 control codes, the space character, and 26 punctuation marks. The set was chosen so that an encoded stream avoids bytes that carry syntactic meaning in common data and command formats. Only six punctuation marks remain usable as trail bytes: !, (, ), ^, _ and ~.
The survivors are exactly the alphanumeric characters plus those six: 10 digits, 52 letters and 6 punctuation marks, 68 in total. This leaves 68 usable page indices, giving the two-byte tier a capacity of 128 × 68 = 8,704 characters.
Addressing
[edit]In a two-byte sequence the trail byte selects a page of 128 characters and the lead byte selects a position, or "slot", within it. This is the reverse of UTF-8, where the leading byte carries the identifying information.
For example, the Devanagari letter na (U+0928) is encoded as the byte pair A8 49:
A8=1 0101000— high bit set, slot 28 hex49=0 1001001— high bit clear, page index 49 hex
Page 49 hex corresponds to the Unicode block beginning at U+0900, so the sequence resolves to U+0900 + 28 hex = U+0928. Holding the lead byte constant and changing only the trail byte selects the same slot in a different script. Conversely, because a script occupies a small fixed set of pages, every character of a given script terminates on the same small set of bytes: Devanagari on I, Thai on R, Korean on six values.
Canonical encoding
[edit]The three-byte tier enumerates every code point from U+0000 upward, including those already reachable in one or two bytes, so each of the 128 ASCII characters and each of the 8,704 two-byte characters can also be expressed in three bytes. A decoder accepting both forms would reproduce UTF-8's overlong encodings: the sequence 80 80 67 would decode to /, which could create security problems for software that validates byte sequences without first enforcing canonical representation. Decoders are therefore required to reject non-canonical sequences, as UTF-8 decoders have been since RFC 3629. The rule is that a three-byte sequence must decode to a code point having no shorter representation, which rejects 8,832 sequences and retains the 1,105,280 that encode characters reachable no other way. The encoder never emits a non-canonical form, and all assigned code points round-trip.
Character allocation
[edit]Pages are of two kinds. The 38 block pages map a 128-code-point-aligned slice of Unicode, so the slot is computed arithmetically. The 30 indexed pages hold a curated list of 128 entries, so the slot requires a table lookup. Chinese and Korean use indexed pages so that characters can be assigned according to the selected repertoire rather than Unicode code-point order, which arranges Han characters by radical and Hangul syllables by jamo composition.
Page map
[edit]The following table lists every trail-byte value. The row gives the high nibble and the column the low nibble; the small character in each cell is the ASCII character that byte represents, and the value beneath it is the page selected. WLD, CHN and KOR mark indexed pages, which hold a curated list rather than an aligned Unicode block.
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | The 32 control codes of rows 0 and 1 can never be a page index, nor can DEL at 7F. | |||||||||||||||
| 1 | ||||||||||||||||
| 2 | SP | ! CHN |
" | # | $ | % | & | ' | ( CHN |
) CHN |
* | + | , | - | . | / |
| 3 | 0 CHN |
1 CHN |
2 CHN |
3 CHN |
4 CHN |
5 CHN |
6 CHN |
7 CHN |
8 300 |
9 308 |
: | ; | < | = | > | ? |
| 4 | @ | A 008 |
B 010 |
C 038 |
D 040 |
E 050 |
F 058 |
G 060 |
H 068 |
I 090 |
J 098 |
K 0A0 |
L 0A8 |
M 0B0 |
N 0C0 |
O 0C8 |
| 5 | P 0D0 |
Q 0D8 |
R 0E0 |
S 0E8 |
T 0F0 |
U 100 |
V 120 |
W 128 |
X 130 |
Y 178 |
Z 1E8 |
[ | \ | ] | ^ WLD |
_ WLD |
| 6 | ` | a 4E0 |
b 4E8 |
c 538 |
d 540 |
e 5B8 |
f 628 |
g 7E8 |
h 8B8 |
i 8D0 |
j 8F8 |
k KOR |
l KOR |
m KOR |
n KOR |
o KOR |
| 7 | p CHN |
q CHN |
r CHN |
s CHN |
t CHN |
u CHN |
v CHN |
w CHN |
x CHN |
y CHN |
z CHN |
{ | | | } | ~ KOR |
DEL |
| Colour | Category | Pages | Indexed pages |
|---|---|---|---|
| World scripts | 30 | 2 WLD | |
| Chinese | 32 | 3+8+11 CHN | |
| Korean | 6 | 5+1 KOR | |
| Excluded | 60 |
Devanagari occupies row 4, column 9 — the ASCII letter I — which is why text in that script produces byte streams containing that letter at every second position.
The page values in the table above are Unicode block addresses with the final hexadecimal digit removed; page 090 therefore denotes the 128 code points beginning at U+0900. The trail byte column gives the ASCII character that terminates every character of that script: Thai always ends on R, Devanagari on I. The scripts they correspond to are:
| Page | Trail byte | Script | Unicode range | Coverage |
|---|---|---|---|---|
008 | A | Latin-1 Supplement | U+0080–U+00FF | 128-code-point slice |
010 | B | Latin Extended-A | U+0100–U+017F | 128-code-point slice |
038 | C | Greek and Coptic | U+0380–U+03FF | 128-code-point slice |
040 | D | Cyrillic | U+0400–U+047F | 128-code-point slice |
050 | E | Armenian and Cyrillic Supplement | U+0500–U+057F | 128-code-point slice |
058 | F | Hebrew and the Armenian tail | U+0580–U+05FF | 128-code-point slice |
060 | G | Arabic | U+0600–U+067F | 128-code-point slice |
068 | H | Arabic, extended letters | U+0680–U+06FF | 128-code-point slice |
090 | I | Devanagari | U+0900–U+097F | 128-code-point slice |
098 | J | Bengali | U+0980–U+09FF | 128-code-point slice |
0A0 | K | Gurmukhi | U+0A00–U+0A7F | 128-code-point slice |
0A8 | L | Gujarati | U+0A80–U+0AFF | 128-code-point slice |
0B0 | M | Odia | U+0B00–U+0B7F | 128-code-point slice |
idx | ^ | Tamil | U+0B80–U+0BFF | Indexed: the 61 assigned characters, the rest of the block being reserved |
0C0 | N | Telugu | U+0C00–U+0C7F | 128-code-point slice |
0C8 | O | Kannada | U+0C80–U+0CFF | 128-code-point slice |
0D0 | P | Malayalam | U+0D00–U+0D7F | 128-code-point slice |
0D8 | Q | Sinhala | U+0D80–U+0DFF | 128-code-point slice |
0E0 | R | Thai | U+0E00–U+0E7F | 128-code-point slice |
0E8 | S | Lao | U+0E80–U+0EFF | 128-code-point slice |
0F0 | T | Tibetan | U+0F00–U+0FFF | First half of the Tibetan block |
100 | U | Myanmar | U+1000–U+109F | First 128 code points of the Myanmar block |
idx | ^ | Georgian | U+10A0–U+10FF | Indexed, on a shared world page (Tamil 61 + Georgian 33 + Mongolian 34 = 128) |
120 | V–X | Ethiopic | U+1200–U+137F | Three pages, 384 characters |
178 | Y | Khmer | U+1780–U+17FF | 128-code-point slice |
idx | ^ | Mongolian | U+1800–U+18AF | Indexed: 34 selected characters of the traditional script |
1E8 | Z | Vietnamese | U+1E00–U+1EFF | Upper half of Latin Extended Additional; Vietnamese shares it with Yoruba, Igbo and Welsh |
300 | 8–9 | Japanese | U+3000–U+30FF | Two pages, 256 characters: CJK punctuation and kana complete |
4E0 | a–j !()0–7 | Chinese | U+4E00–U+9FFF | 32 pages in total: 10 block pages covering 1,280 characters as aligned slices, and 22 indexed pages holding the 2,816 most frequent; simplified only |
AC0 | k–o ~ | Korean | U+AC00–U+D7A3 | Indexed 6 pages: 768 most frequent syllables of 11,172 |
| Page range | UTF-8 | UCE-8 | Trail bytes | Pages |
|---|---|---|---|---|
008–068 | 2 bytes | 2 bytes | A–H | 8 |
090 onward | 3 bytes | 2 bytes | I–Z, a–z, 0–9 and !()^_~ | 60 |
The shading marks the boundary the encoding is built around. The pages up to 068 hold scripts that already cost two bytes in UTF-8, so UCE-8 holds them at the same size rather than reducing it. The pages from 090 onward hold scripts whose characters generally require three bytes in UTF-8, and it is those the saving applies to.
Character tables
[edit]The 30 indexed pages are backed by three ordered strings, one per category. A character's slot is its position in the string, so the tables are the encoding: reordering them would change the meaning of every byte sequence already written. They are consequently frozen.
| Table | Pages | Characters | Contents |
|---|---|---|---|
CHINESE_CHARS | 22 | 2,816 | Han characters by frequency, excluding those the ten Chinese block pages already reach |
| 十卜八入儿匕几刁刀力干工土士才大小山巾千川夕勺凡广门尸己 … | |||
KOREAN_CHARS | 6 | 768 | Hangul syllables ordered by the frequency of their component jamo |
| 아이가으나기사하다오니그시히디안어우느자마스고흐드알인바 … | |||
WORLD_CHARS | 2 | 256 | Tamil, Georgian and Mongolian on one page; Cyrillic extensions, additional Latin letters, currency signs and fullwidth punctuation on the other |
| ஂஃஅஆஇஈஉஊஎஏஐஒஓஔகஙசஜஞடணதநனபமயர … | |||
| Total | 30 | 3,840 | |
Each table fills its pages exactly, 30 × 128 = 3,840 entries with none unused. The complete tables are too long to reproduce here; they appear in full in the reference implementation.[1]
Comparison with other encodings
[edit]UCE-8 is not the first encoding to represent East Asian scripts in two bytes. Shift JIS, EUC-KR and Big5 all predate it and encode their respective scripts in two bytes, but each covers a single language community and they are mutually incompatible. GB 18030, which supersedes GBK, likewise encodes Chinese in two bytes but extends to the whole of Unicode through a four-byte form, at the cost of four bytes for scripts outside its two-byte repertoire. UTF-16 encodes the entire Basic Multilingual Plane in two bytes but requires two bytes for ASCII as well.
| Script | UTF-8 | UTF-16 | GB 18030 | UCE-8 |
|---|---|---|---|---|
| ASCII | 1 | 2 | 1 | 1 |
| Cyrillic, Greek | 2 | 2 | 2 | 2 |
| Hebrew, Arabic | 2 | 2 | 4 | 2 |
| Indic, Thai, Ethiopic | 3 | 2 | 4 | 2 |
| Chinese, Japanese kana | 3 | 2 | 2 | 2 |
| Korean | 3 | 2 | 4 | 2 |
| Supplementary planes | 4 | 4 | 4 | 3 |
Unicode also defines two compression schemes with related goals, the Standard Compression Scheme for Unicode (UTS #6) and BOCU-1 (UTS #12), neither of which is widely deployed.
Detection and compatibility
[edit]Coexistence with other encodings
[edit]UCE-8 and UTF-8 are close to mutually exclusive: a stream well formed in one is, with rare exceptions, malformed in the other. A reader can therefore establish which of the two a buffer uses by attempting both decodings rather than consulting a label, an extension or a byte order mark. Plain ASCII satisfies both and decodes identically either way, so no distinction is required for it.
The specified rule gives UTF-8 precedence: a buffer that decodes as valid UTF-8 is read as UTF-8, and only a buffer UTF-8 rejects is read as UCE-8. Existing UTF-8 files are consequently never reinterpreted.
Buffers valid in both encodings and decoding differently can be constructed, but they are few and enumerable. Their number is exactly 130,560, being the 30 valid UTF-8 two-byte lead bytes (C2–DF) multiplied by the 64 continuation bytes (80–BF) and the 68 permitted UCE-8 terminators; no other UTF-8 form can collide, since three- and four-byte characters end in a continuation byte that cannot terminate a UCE-8 sequence. Read as UCE-8, every one of them yields a code point between U+8C400 and U+CAEFF, in planes 8 to 12, which Unicode has never assigned, so a decoder that rejects unassigned results rejects the whole set.
The same approach extends beyond UTF-8. A reader that attempts each candidate decoding and assesses the result can also distinguish UCE-8 from GB 18030, GBK, Big5, Shift JIS and EUC-KR: a stream in any of those is rejected by a UCE-8 decoder, on the byte grammar or on the plane test, and each legacy encoding produces text in its own script when read correctly and characters outside it when read wrongly.
Overlap with the GB encodings and its resolution
[edit]GB 18030, the Chinese national standard, and its subset GBK use continuation-byte ranges that largely coincide with UCE-8's terminating bytes: 65 of the 68 are also valid GB continuation bytes. The overlap cannot be designed away, but UCE-8 text produces little ambiguity in practice. UCE-8 text presented to readers that have not been updated for UCE-8 rarely forms valid GB, because eleven of the 32 pages carrying Chinese are addressed by bytes that GB accepts as no ordinary trail byte: eight are digits, which GB 18030 admits only in its four-byte form, and three are the punctuation marks !, ( and ), which fall below GB's trail range altogether. A single such byte invalidates the whole sequence containing it, so the likelihood of a valid GB reading falls as the text lengthens.
A reader updated for UCE-8 is not exposed to that risk, because it tests the buffer instead of assuming an encoding. Valid UTF-8 is read as UTF-8 outright; everything else passes through five stages, each removing part of what the one before it admits.
- The byte grammar. A UCE-8 character is at most three bytes, of which only the last has its high bit clear. Consecutive GB characters place a high byte where UCE-8 requires a terminator, so ordinary GB prose, in which Han characters follow one another, fails here before anything is decoded at all.
- The page-index requirement. The terminating byte must be one of the 68 permitted page indices. This refuses a further 1,008 of GB 18030's 23,940 two-byte characters — 126 for each of
@,[,\,],`,{,|and}, the eight values in GB's trail range that are among the 60 UCE-8 excludes. The exclusion list chosen for stream safety does detection work as a side effect. - The shortest-form requirement. A three-byte sequence must decode to a code point that has no shorter representation. A sequence that clears the grammar but spells a character already reachable in one or two bytes is non-canonical and is refused, which removes a further part of what remains.
- The plane test. The decoded result must not fall in planes 6 to 15 — wider than UTF-8 alone would need, its ambiguous sequences occupying planes 8 to 12 only.
- Assessment of the decoded text. What the first four stages admit is judged as text rather than as bytes. A reading coherent in no script is refused, and the buffer is then scored against the legacy encodings to establish which of them it is.
Limitations
[edit]- Software support. Existing software cannot read UCE-8 until it is modified to do so; to an unmodified reader a UCE-8 file appears to be malformed UTF-8. The limitation is one of deployment rather than of the format itself, since UCE-8 and UTF-8 can coexist in the same directory or data stream without labelling, provided each reader implements the detection rule described above. Detection nonetheless requires the complete buffer in advance, and must be implemented by each reader, so it does not assist software that is unaware of the encoding — grep, a database column or a terminal that has never been told the encoding exists remain unaffected by it. The specification restricts its scope to contemporary text and does not require implementations to decode legacy single-byte encodings, a restriction with precedent in JSON, which requires UTF-8, and in HTML5, which excludes legacy encodings from conformance.
- Byte order does not follow code-point order. Because the page index occupies the second byte, a byte-wise comparison orders by slot before page; the discrepancy therefore arises even between two block pages, and is compounded within indexed pages, which are ordered by frequency rather than by code point. ASCII text is unaffected. Decoding the text and sorting the resulting code points restores the expected order; what is not available is binary search, range querying or index ordering over the encoded bytes without decoding them, operations for which UTF-8's byte order does correspond to code-point order.
- Context-dependent bytes. A trail byte cannot be distinguished from a standalone ASCII character by inspection alone; determining which requires examining the preceding byte. UTF-8, by contrast, marks continuation bytes with a reserved
10xxxxxxprefix. The omission is a deliberate trade: reserving such a prefix would reduce the two-byte tier from 8,704 characters to 4,096. - Substring search. Because trail bytes are ordinary ASCII digits and letters, a naive byte-level substring search can report false positives. These can be eliminated by testing the byte preceding each candidate match, described under Character boundaries below. The test applies only to code that performs the search, so tools that are unaware of the encoding — grep, SQL
LIKE, or a general-purpose search index — remain affected. - Error resilience. UTF-8 is a self-synchronizing code, so a corrupted byte damages only the character containing it. UCE-8 is not: a terminating byte whose high bit is corrupted ceases to terminate, and its character runs on into the next. The damage is nonetheless bounded, since a character is at most three bytes and a decoder meeting three consecutive lead bytes resumes at the next terminating byte. Nor does it spare ASCII, as most of the limitations above do: a flipped high bit turns a standalone ASCII byte into a lead byte that swallows its neighbour. GB 18030 has no such ceiling: its second and fourth bytes take ASCII values. A single corrupted bit costs UTF-8 one character, UCE-8 two, and GB 18030 the rest of the text.
- Fixed repertoire. Reordering the curated character tables would alter the meaning of the 3,840 two-byte sequences that address indexed pages, invalidating previously encoded Chinese, Korean, Tamil, Georgian and Mongolian text; ASCII, the block pages and the three-byte tier are arithmetic and would be unaffected. The resulting permanence applies equally to correct and incorrect frequency choices.
- Interaction with compression. Much of the size advantage over UTF-8 is reduced when general-purpose compression such as gzip or Brotli is applied, since UTF-8's repeated lead bytes are themselves highly compressible: a compressor removes by observation much of what UCE-8 removes by design. On short passages the encoded form may be marginally larger than UTF-8, as Brotli's built-in dictionary consists largely of UTF-8 text that UCE-8 output cannot match.
Implementation
[edit]A reference implementation in Python, together with the frozen character tables and a browser-based demonstration, is published in a public repository.[1] A second implementation in C (uce.c) is provided in the same repository; it holds the curated tables as flat arrays and resolves the indexed pages through a perfect hash.[1]
Encoding
[edit]The encoder tries the mappings in order of cost. A code point below U+0080 passes straight through as a single byte. Otherwise the block pages are consulted first, since their slot is arithmetic: the page is the code point's aligned 128-point slice, and the slot is its position within that slice. If no block page holds it, the three curated tables are searched in turn, and the slot is the character's position in the table. Anything still unmatched falls to the three-byte tier.
That tier has no page table left to consult, so it distributes the remaining code points arithmetically: the remainder modulo 68 selects one of the legal terminating bytes and the quotient becomes the two lead bytes. This is why the tier's capacity is exactly 128 × 128 × 68, and why the encoder cannot emit an illegal terminator even in its fallback path.
Decoding
[edit]Decoding reverses the same order. A byte below 0x80 is an ASCII character on its own. Otherwise the high bit of the second byte selects the tier: if it is clear, the sequence is two bytes and that second byte is the page index, resolved either arithmetically for a block page or through the corresponding table for an indexed one. If it is set, a third byte follows and closes a three-byte sequence, whose code point is the two lead bytes multiplied by 68 plus the terminator's index. A three-byte sequence is then rejected if the code point it yields has a shorter representation.
Character boundaries
[edit]Because a trail byte is indistinguishable from a standalone ASCII character in isolation, a byte-level substring search over a UCE-8 stream can match the interior of a multi-byte character. The effect is pronounced for scripts whose page index happens to be a common ASCII character: Devanagari's page index is 0x49, the letter I, so Devanagari text contains that byte at every second position. Searching such a stream for the letter I returns one match per character in addition to any genuine occurrence of that letter.
The ambiguity is resolved by a single test on the preceding byte. A low byte always terminates a character, so a position begins a character precisely when the byte before it is also low, or when it is the start of the stream.
Only the first byte of a candidate match needs this test. If that byte begins a character and is itself low, it also terminates one, so the following byte begins a character in turn, and the property carries along the whole search term. The same test governs where a stream may be split: a cut is safe unless the byte before it is a lead byte.
See also
[edit]References
[edit]- 1 2 3 "UCE-8 reference implementation". GitLab. Retrieved 2026-08-11.