Reference articles on history, science, culture and more
Encyclopedia

Cork encoding

Latin script character encoding used by LaTeX

The Cork (also known as T1 or EC) encoding is a character encoding used for encoding glyphs in fonts. It is named after the city of Cork in Ireland, where during a TeX Users Group (TUG) conference in 1990 a new encoding was introduced for LaTeX. It contains 256 characters supporting most west- and east-European languages with the Latin alphabet.

01Details

In 8-bit TeX engines the font encoding has to match the encoding of hyphenation patterns where this encoding is most commonly used. In LaTeX one can switch to this encoding with \usepackage[T1]{fontenc}, while in ConTeXt MkII this is the default encoding already. In modern engines such as XeTeX and LuaTeX Unicode is fully supported and the 8-bit font encodings are obsolete.

02Character set

Cork encoding
0 1 2 3 4 5 6 7 8 9 A B C D E F
0x `0060 ´00B4 ˆ02C6 ˜02DC ¨00A8 ˝02DD ˚02DA ˇ02C7 ˘02D8 ¯00AF ˙02D9 ¸00B8 ˛02DB 201A 2039 203A
1x 201C 201D 201E «00AB »00BB , 2013 , 2014 ZWSP
200B
2080 ı0131 ȷ0237 FB00 FB01 FB02 FB03 FB04
2x ␣2423 ! " # $ % & 2019 ( ) * + , - . /
3x 0 1 2 3 4 5 6 7 8 9 : ; < = > ?
4x @ A B C D E F G H I J K L M N O
5x P Q R S T U V W X Y Z [ \ ] ^ _
6x 2018 a b c d e f g h i j k l m n o
7x p q r s t u v w x y z { | } ~ SHY
8x Ă0102 Ą0104 Ć0106 Č010C Ď010E Ě011A Ę0118 Ğ011E Ĺ0139 Ľ013D Ł0141 Ń0143 Ň0147 Ŋ014A Ő0150 Ŕ0154
9x Ř0158 Ś015A Š0160 Ş015E Ť0164 Ţ0162 Ű0170 Ů016E Ÿ0178 Ź0179 Ž017D Ż017B IJ0132 İ0130 đ0111 §00A7
Ax ă0103 ą0105 ć0107 č010D ď010F ě011B ę0119 ğ011F ĺ013A ľ013E ł0142 ń0144 ň0148 ŋ014B ő0151 ŕ0155
Bx ř0159 ś015B š0161 ş015F ť0165 ţ0163 ű0171 ů016F ÿ00FF ź017A ž017E ż017C ij0133 ¡00A1 ¿00BF £00A3
Cx À Á Â Ã Ä Å Æ Ç È É Ê Ë Ì Í Î Ï
Dx Ð Ñ Ò Ó Ô Õ Ö Œ0152 Ø Ù Ú Û Ü Ý Þ SS1E9E
Ex à á â ã ä å æ ç è é ê ë ì í î ï
Fx ð ñ ò ó ô õ ö œ0153 ø ù ú û ü ý þ ß00DF

03Supported languages

The encoding supports most European languages written in Latin alphabet. Notable exceptions are:

Languages with slightly suboptimal support include:

  • Galician language, Portuguese language and Spanish language, due to the lack of characters ª and º, which are not superscript versions of lowercase "a" and "o" (superscripts are thinner) and they are often underlined
  • Croatian language, Bosnian language, Serbian language, due to the shared use of the slot for Đ
  • Turkish language, due to dotless i having different uppercase and lowercase combinations than in other languages
  • Romanian language, due to the characters "Ş ş Ţ ţ" (with a cedilla) being typographically considered incorrect by modern standards, with the expected correct forms being "Ș ș Ț ț" (with a comma below) - though when the encoding was developed, it was arguably considered acceptable at that time, but the status of support retroactively changed to suboptimal or insufficient when the Unicode codepoints were disunified.
Watch videos about Cork encodingExplainers and documentaries on YouTube (opens in a new tab)

Sources and credits

This article is adapted from the Wikipedia article Cork encoding, written by its contributors and licensed under CC BY-SA 4.0. Fathomly has changed the layout, removed citation markers, navigation and maintenance notices, and adjusted punctuation. This adapted version is shared under the same license. For references, see the original article.

Fathomly is not affiliated with or endorsed by the Wikimedia Foundation. Spotted a problem? Tell us.