Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a450900edd51b071

Jump to content

Draft:TeX font encodings

From Wikipedia, the free encyclopedia


(La)TeX font encodings are character-to-glyph mapping schemes used by the TeX typesetting system and related software such as LaTeX. Unlike modern Unicode encodings, TeX font encodings traditionally assign characters to positions within an 8-bit font, where each slot corresponds to a glyph rather than a universal character identity. These encodings were designed primarily for high-quality typesetting rather than text interchange.[1]

History

[edit]

Early TeX systems commonly used 7-bit fonts that contain at most 128 glyph slots. Because of this limitation, four distinct font encodings were developed for different purposes: one for text and three for various math symbols (in modern TeX language these are called OT1, OML, OMS, and OMX, but back then they did not have distinct names).

Because 7 bits is way too small to fit an entire multilingual alphabet, the alphabet was simplified into base letters and separate diacritics which had their own codepoints. Donald Knuth, the creator of TeX was initially satisfied but later, Knuth realized that 128 different characters for the text input were not enough to accommodate foreign languages (his solution for diacritics turned out to be inefficient and problematic as it caused problems with kerning, hyphenation, and text copying). To fix that, when version 3.0 of TeX was released, it gained the ability to work with 8-bit inputs, allowing 256 different characters in the text input. [2]

Since support for 8-bit encodings was added, the natural question was: What should the ideal multilingual 8-bit encoding look like? To address this issue, a team at the TUG Annual General Meeting held in Cork, Ireland, defined a 256-glyph font encoding, which includes accented letters and non-ASCII characters essential for representing the majority of Western European languages (and a few Eastern European ones) without needing commands such as \accent that create characters from diacritic parts. This “Cork” encoding has been implemented in several fonts created with Metafont, as well as in various virtual-font mappings of different font series. Since the meeting, considerable effort has gone into designing encodings for numerous other text fonts compatible with TeX, with the Cork encoding (T1) shaping the creation of many such encodings utilized in various fonts globally.[1]

Differences

[edit]

Traditional TeX font encodings differ from general-purpose text encodings such as ASCII or Unicode because they were designed for typesetting specific fonts rather than representing plain text. Here are the most important differences:

  • In normal encodings, characters 0x20 and 0xA0 are often whitespace characters, and 0x00-0x1F and 0x7F-0x9F are control characters. In TeX, the role of control characters and whitespace characters is taken over by TeX's formatting itself, meaning that TeX does not need these characters. This means that in TeX font encodings, all slots are printable, meaningful characters. 0x20 isn't space, 0x00 isn't control character NUL, and 0xA0 is not a non-breaking space. This discrepancy has caused problems when early PDF readers were developed. For example, the character 0x00 of TeX encodings (OT1 Greek Gamma) was known to break old PDF readers because the PDF reader, unfamiliar with TeX encodings, confused it with the ASCII code 0x00.[3]
  • In other text encodings, such as Unicode or Code page 437, the set of all characters that encoding supports is all that can be typed. In TeX encodings, some characters are instead simulated by macros from other characters[4], meaning it is possible to "type" characters that don't exist by combining other characters. [5][6][7][8][9][10][11][12]

128+ glyph encodings (text)

[edit]

OT1

[edit]

OT1 (aka TeX text) is a 7-bit TeX encoding developed by Donald E. Knuth.[13][1]


OT1[1]
0 1 2 3 4 5 6 7 8 9 A B C D E F
0x Γ Δ Θ Λ Ξ Π Σ Υ Φ Ψ Ω ff fi fl ffi ffl
1x ı ȷ ` ´ ˇ ˘ ˉ ˚ ¸ ß æ œ ø Æ Œ Ø
2x ̷ ! ” # $ % & ’ ( ) * + , - . /
3x 0 1 2 3 4 5 6 7 8 9 : ; ¡ = ¿ ?
4x @ A B C D E F G H I J K L M N O
5x P Q R S T U V W X Y Z [ “ ] ˆ ˙
6x ‘ a b c d e f g h i j k l m n o
7x p q r s t u v w x y z – — ˝ ˜ ¨

Notes

[edit]
  • The letters A, B, E, H, I, K, M, N, O, P, T, X, Z are used both as the Latin letters and their Greek homoglyphs. The remaining Greek letters are found in positions 0x00-0x0A.
  • The diacritics (0x12-0x18, 0x20, 0x5E-0x5F, 0x7D-0x7F) are used as combining characters to compose different letters with diacritics. 0x20 is used to create the letter Ł[1] , but can be reused for other letters with slanted strokes (which are not even in Unicode but may be needed for nonstandard use[14]).

OT2

[edit]

OT2 (aka UW cyrillic encoding[1]) is a 7-bit TeX encoding developed for typesetting text in the Cyrillic script.[1] It predates the later T2A, T2B, and T2C encodings.

OT2[1]
0 1 2 3 4 5 6 7 8 9 A B C D E F
0x Њ Љ Џ Э І Є Ђ Ћ њ љ џ э і є ђ ћ
1x Ю Ж Й Ё Ѵ Ѳ Ѕ Я ю ж й ё ѵ ѳ ѕ я
2x ¨ ! ” Ѣ ˘[a] % ´ ’ ( ) * ѣ , - . /
3x 0 1 2 3 4 5 6 7 8 9 : ; « ı[b] » ?
4x ˘[a] А Б Ц Д Е Ф Г Х И Ј К Л М Н О
5x П Ч Р С Т У В Щ Ш Ы З [ “ ] Ь Ъ
6x ‘ а б ц д е ф г х и ј к л м н о
7x п ч р с т у в щ ш ы з – — № ь ъ

Notes

[edit]
  1. 1 2 OT2 has a distinction between two characters that are not distinguished in Unicode: the wide breve 0x24 and the smaller breve 0x40.[15]
  2. ↑ This is a dotless version of і, used to compose diacritic variants like ї.

OT3

[edit]

OT3 (aka UW IPA encoding[1]) is a 7-bit TeX font encoding intended for IPA. It was never really used with LaTeX 2ε following the release of the TIPA package which offers much better support for IPA and unlike OT3, supports other linguistic notations such as Germanic and Indo-European transcriptions, as well as a large portion of extIPA. No ot3enc.def file was ever produced.[1] This encoding is supported by the wsuipa package.[16]

OT3[1][17]
0 1 2 3 4 5 6 7 8 9 A B C D E F
0x ɐ ɑ α ɒ ʌ ƀ b[a] bⳆ[b] ɓ β ȼ ɕ ʗ đ d[c] dⳆ[d]
1x ɗ ɖ ʤ ð ᴅ ə ɚ ɘ ɛ ɜ ɝ ɞ ɡ ɠ ɢ γ
2x ɣ ɤ ƕ ħ ɦ ɧ ɥ ɨ ᵻ[e] ɩ ɪ ᵻ[e] ɟ ɫ ƚ ɬ
3x ɭ ɮ ꟛ ƛ ɱ ɯ ɰ ɲ ŋ ɳ ɴ ʘ ɵ ɔ ꞷ ɷ
4x ꝏ ᵽ þ ɸ ɾ ɼ ɽ ɹ ɻ ɺ ʀ ʁ ʂ ʃ ʆ σ
5x ʈ ʧ ʇ θ ʉ ꞹ ʊ ᴜ ᵾ ʋ ʍ χ ʎ ʏ ʑ ʐ
6x ʒ ʓ ʔ ʕ ʖ ɂ ꟏ ̪ ˈ ˌ ̩ ̚ ˔ ˕ ̘ ̙
7x ˑ ː ̤ ʽ ˄ ˅ ˂ ˃ ˚ ˳ ̜ ̴ ˷ ˬ ˛ ̯

Notes

[edit]
  1. ↑ b with a stroke through the bowl. Not supported by Unicode.
  2. ↑ slashed b. Not supported by Unicode.
  3. ↑ d with a stroke through the bowl. Not supported by Unicode.
  4. ↑ slashed d. Not supported by Unicode.
  5. 1 2 The characters 0x28 and 0x2B are distinct characters but are not distinguished by Unicode. The former is a barred dotless i and the latter is a barred small capital i.

OT4

[edit]

OT4 (also known as Polish TeX encoding) is an 8-bit TeX font encoding designed for the Polish language. The lower 128 bytes are identical to OT1. The OT4 encoding was created because OT1 was missing the ogonek diacritic, even though the Ł was included.[1]

OT1[1]
0 1 2 3 4 5 6 7 8 9 A B C D E F
0x Γ Δ Θ Λ Ξ Π Σ Υ Φ Ψ Ω ff fi fl ffi ffl
1x ı ȷ ` ´ ˇ ˘ ˉ ˚ ¸ ß æ œ ø Æ Œ Ø
2x ̷ ! ” # $ % & ’ ( ) * + , - . /
3x 0 1 2 3 4 5 6 7 8 9 : ; ¡ = ¿ ?
4x @ A B C D E F G H I J K L M N O
5x P Q R S T U V W X Y Z [ “ ] ˆ ˙
6x ‘ a b c d e f g h i j k l m n o
7x p q r s t u v w x y z – — ˝ ˜ ¨
8x   Ą Ć       Ę       Ł Ń        
9x   Ś               Ź   Ż        
Ax   ą ć       ę       ł ń     « »
Bx   ś               ź   ż        
Cx                                
Dx       Ó                        
Ex                                
Fx       ó                       „

Notes

[edit]
  • OT4 was created to simplify Polish typesetting in TeX before T1 was available.[1]

OT5

[edit]

The name OT5 is not currently allocated.

OT6

[edit]

OT6 is a TeX font encoding for the Armenian alphabet. It was created by Sergueï Dachian, and was allocated to permit use of his Armenian fonts in a standard LaTeX environment.[1][18]


OT6[1]
0 1 2 3 4 5 6 7 8 9 A B C D E F
0x   ՟ ՙ Ձ Ղ Ճ Է Ը Թ Ժ Շ Չ Ռ Ծ Փ Օ
1x   ֏ ֍ ձ ղ ճ է ը թ ժ շ չ ռ ծ փ օ
2x և ՜ ” # $ % & ՚ * + , - . /
3x 0 1 2 3 4 5 6 7 8 9 ։ ; « = » ՞
4x @ Ա Բ Ց Դ Ե Ֆ Գ Հ Ի Ջ Կ Լ Մ Ն Ո
5x Պ Ք Ր Ս Տ ՈՒ[a] Վ Ւ Խ Յ Զ [ “ ] { }
6x ՝ ա բ ց դ ե ֆ գ հ ի ջ կ լ մ ն ո
7x պ ք ր ս տ ու[a] վ ւ խ յ զ ֊ ՛ — ! ?

Notes

[edit]
  1. 1 2 In OT6, the digraphs ՈՒ and ու are given single codepoints to make Latin-script text convert to Armenian more easily.

256 glyph encodings (text)

[edit]

T1

[edit]

The T1 encoding (aka Cork encoding) is a character encoding used for encoding glyphs in fonts.[19] It is named after the city of Cork in Ireland, where during a TeX Users Group (TUG) conference in 1990 a new encoding was introduced for LaTeX.[19] It contains 256 characters supporting most west- and east-European languages with the Latin alphabet.[20] In 8-bit TeX engines the font encoding has to match the encoding of hyphenation patterns where this encoding is most commonly used.[21] . It is now very common.

Cork encoding
0 1 2 3 4 5 6 7 8 9 A B C D E F
0x `
0060
´
00B4
ˆ
02C6
˜
02DC
¨
00A8
˝
02DD
˚
02DA
ˇ
02C7
˘
02D8
¯
00AF
˙
02D9
¸
00B8
˛
02DB
‚
201A
‹
2039
›
203A
1x “
201C
”
201D
„
201E
«
00AB
»
00BB
–
2013
—
2014
ZWSP[a]
200B
₀[b]
2080
ı[c]
0131
ȷ[c]
0237
ff
FB00
fi
FB01
fl
FB02
ffi
FB03
ffl
FB04
2x ␣
2423
! " # $ % & ’
2019
( ) * + , - . /
3x 0 1 2 3 4 5 6 7 8 9 : ; < = > ?
4x @ A B C D E F G H I J K L M N O
5x P Q R S T U V W X Y Z [ \ ] ^ _
6x ‘
2018
a b c d e f g h i j k l m n o
7x p q r s t u v w x y z { | } ~ SHY[d]
8x Ă
0102
Ą
0104
Ć
0106
Č
010C
Ď
010E
Ě
011A
Ę
0118
Ğ
011E
Ĺ
0139
Ľ
013D
Ł
0141
Ń
0143
Ň
0147
Ŋ
014A
Ő
0150
Ŕ
0154
9x Ř
0158
Ś
015A
Š
0160
Ş
015E
Ť
0164
Ţ
0162
Ű
0170
Ů
016E
Ÿ
0178
Ź
0179
Ž
017D
Ż
017B
IJ
0132
İ
0130
đ
0111
§
00A7
Ax ă
0103
ą
0105
ć
0107
č
010D
ď
010F
ě
011B
ę
0119
ğ
011F
ĺ
013A
ľ
013E
ł
0142
ń
0144
ň
0148
ŋ
014B
ő
0151
ŕ
0155
Bx ř
0159
ś
015B
š
0161
ş
015F
ť
0165
ţ
0163
ű
0171
ů
016F
ÿ
00FF
ź
017A
ž
017E
ż
017C
ij
0133
¡
00A1
¿
00BF
£
00A3
Cx À Á Â Ã Ä Å Æ Ç È É Ê Ë Ì Í Î Ï
Dx Ð[e] Ñ Ò Ó Ô Õ Ö Œ
0152
Ø Ù Ú Û Ü Ý Þ SS[f]
1E9E
Ex à á â ã ä å æ ç è é ê ë ì í î ï
Fx ð ñ ò ó ô õ ö œ
0153
ø ù ú û ü ý þ ß
00DF

Notes

[edit]
  • Hexadecimal values under the characters in the table are the Unicode character codes.
  • The first 12 characters are often used as combining characters.
  1. ↑ 0x17 is dubbed a “compound word mark” (CWM) in the Cork encoding, and is an innovation of this standard. It is an invisible character that separates compounds in a complex word, for instance in German, in order to disallow esthetic ligatures at compound boundaries.[20] It is mapped to the Unicode “zero-width space” (ZWSP, U+200B), defined at about the same time, whose purpose is similar, if not identical.
  2. ↑ 0x18 is a “small o”, used to compose ‰ or ‱ (or arbitrary smaller quantities) out of percent sign (%).[20]
  3. 1 2 Dotless i and dotless j may be used to compose accented variants like i with macron (ī).
  4. ↑ 0x7F is the hyphenation character, not really a soft hyphen (SHY) as defined by Unicode.
  5. ↑ 0xD0 is used both as Eth (Ð, U+00D0) and as D with stroke (Đ, U+0110) which might be a problem at some occasions (like copying text from PDF, hyphenation, ...)
  6. ↑ 0xDF contains an uppercase variant of ß. Traditionally, it was SS (two letters S), but some newer fonts may use ẞ.[22] It allows TeX to automatically convert the German lowercase ß into the uppercase form.

T2A

[edit]

The T2A encoding (aka Cork encoding) is a character encoding used for encoding glyphs in fonts.[19] It is named after the city of Cork in Ireland, where during a TeX Users Group (TUG) conference in 1990 a new encoding was introduced for LaTeX.[19] It contains 256 characters supporting most west- and east-European languages with the Latin alphabet.[20] In 8-bit TeX engines the font encoding has to match the encoding of hyphenation patterns where this encoding is most commonly used.[21]

T2A encoding
0 1 2 3 4 5 6 7 8 9 A B C D E F
0x `
0060
´
00B4
ˆ
02C6
˜
02DC
¨
00A8
˝
02DD
˚
02DA
ˇ
02C7
˘[a]
02D8
¯
00AF
˙
02D9
¸
00B8
˛
02DB
Ӏ
04C0
〈
2329
〉
232A
1x “
201C
”
201D
̑
0311
̏
030F
˘[a]
02D8
–
2013
—
2014
ZWNJ
200C
₀[b]
2080
ı[c]
0131
ȷ[c]
0237
ff
FB00
fi
FB01
fl
FB02
ffi
FB03
ffl
FB04
2x ␣
2423
! " # $ % & ’
2019
( ) * + , - . /
3x 0 1 2 3 4 5 6 7 8 9 : ; < = > ?
4x @ A B C D E F G H I[c] J[c] K L M N O
5x P Q R S T U V W X Y Z [ \ ] ^ _
6x ‘
2018
a b c d e f g h i[c] j[c] k l m n o
7x p q r s t u v w x y z { | } ~ SHY[d]
8x Ґ Ғ Ђ Ћ Һ Җ Ҙ Љ Ї Қ Ҡ Ҝ Ӕ Ң Ҥ Ѕ
9x Ө Ҫ Ў Ү Ұ Ҳ Џ Ҹ Ҷ Є Ә Њ Ё № ¤ §
Ax ґ ғ ђ ћ һ җ ҙ љ ї қ ҡ ҝ ӕ ң ҥ ѕ
Bx ө ҫ ў ү ұ ҳ џ ҹ ҷ є ә њ ё „ « »
Cx А Б В Г Д Е Ж З И Й К Л М Н О П
Dx Р С Т У Ф Х Ц Ч Ш Щ Ъ Ы Ь Э Ю Я
Ex а б в г д е ж з и й к л м н о п
Fx р с т у ф х ц ч ш щ ъ ы ь э ю я

Notes

[edit]
  1. 1 2 T2A has a distinction between two characters that are not distinguished in Unicode: the wide breve 0x14 and the smaller breve 0x08.
  2. ↑ 0x18 is a “small o”, used to compose ‰ or ‱ (or arbitrary smaller quantities) out of percent sign (%).
  3. 1 2 3 4 5 6 The characters I, J, i, j, ı, and ȷ are both used as the Latin letters I and J, as well as the Cyrillic letters І and Ј.[23]
  4. ↑ 0x7F is the hyphenation character, not really a soft hyphen (SHY) as defined by Unicode.


References

[edit]
  1. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 Mittelbach, Frank; Fairbairns, Robin; Lemberg, Werner (2016-02-18) [1995]. "LaTeX font encodings" (PDF). LATEX3 Project Team. Archived (PDF) from the original on 2017-07-10. Retrieved 2017-07-10.
  2. ↑ Hoenig, Alan (1998). TeX Unbound: LaTeX & TeX Strategies for Fonts, Graphics, & More. Oxford University Press. ISBN 978-0-19-509686-6.
  3. ↑ https://tex.stackexchange.com/a/761663/423378
  4. ↑ https://texdoc.org/serve/tipaman.pdf/0
  5. ↑ https://tex.stackexchange.com/questions/612457/why-does-making-a-d-with-a-bar-through-it-not-behave-like-an-h-with-a-bar-throug
  6. ↑ https://tex.stackexchange.com/questions/486049/how-to-create-a-command-for-the-strange-m-symbol-in-latex
  7. ↑ https://tex.stackexchange.com/questions/472434/how-do-i-italicize-a-crossed-h
  8. ↑ https://tex.stackexchange.com/questions/527065/how-to-typeset-upright-%c4%a7
  9. ↑ https://tex.stackexchange.com/questions/460110/h-character-with-crossbeam
  10. ↑ https://tex.stackexchange.com/questions/633116/how-to-write-an-e-with-a-circle-around-the-middle-line
  11. ↑ https://tex.stackexchange.com/a/711888/423378
  12. ↑ https://tex.stackexchange.com/a/238491/423378
  13. ↑ Knuth, Donald E. (May 1989). The TEXbook (PDF). Computers & Typesetting. Vol. A (Eight printing ed.). p. 427. Archived from the original (PDF) on 2004-09-24. Retrieved 2020-05-01.
  14. ↑ https://tex.stackexchange.com/a/663838/423378
  15. ↑ https://ctan.math.washington.edu/tex-archive/macros/latex/required/cyrillic/ot2.pdf
  16. ↑ https://ctan.org/pkg/wsuipa
  17. ↑ https://www.mn.uio.no/ifi/tjenester/it/hjelp/latex/ipaman.pdf
  18. ↑ https://ctan.gust.org.pl/tex-archive/language/armenian/armtex/manual-e.pdf
  19. 1 2 3 4 Petrlik, Lukas (1996-06-19). "The Czech and Slovak Character Encoding Mess Explained". cs-encodings-faq. 1.10. Archived from the original on 2016-06-21. Retrieved 2016-06-21.
  20. 1 2 3 4 Ferguson, Michael (1990), "Report on Multilingual Activities" (PDF), TUGboat, 11 (4): 514–516
  21. 1 2 TeX hyphenation patterns
  22. ↑ https://tex.stackexchange.com/questions/754903/how-to-make-the-letter-%e1%ba%9e-in-latex
  23. ↑ https://tex.stackexchange.com/questions/578452/ukrainian-%D1%96-replaced-with-latin-on-compilation