Reference articles on history, science, culture and more
Encyclopedia

CJK Unified Ideographs

Encoding for shared Han characters

The Chinese, Japanese and Korean (also known as CJK) scripts share a common background, collectively known as CJK characters. During the process called Han unification, the common (shared) characters were identified and named CJK Unified Ideographs. As of Unicode 17.0, Unicode defines a total of 101,996 characters.

The term ideographs is a misnomer, as the Chinese script is not ideographic but rather logographic, but was chosen for being more common in English.

Until the early 20th century, Vietnam also used Chinese characters (Chữ Nôm), so sometimes the abbreviation CJKV is used.

01Ordering

The ordering of CJK Unified Ideographs within Unicode blocks (not counting those added to the block later) was initially determined by consulting the following four dictionaries. Primarily, they were arranged in Kangxi Dictionary order, with the other dictionaries consulted, in order, for characters not found in the Kangxi Dictionary, to determine which Kangxi Dictionary character they should follow in the ordering.

  1. Kangxi Dictionary
  2. Dai Kan-Wa Jiten
  3. Hanyu Da Zidian
  4. Dae Jaweon

This system is not used for more recently-added Unicode blocks. The Ideographic Research Group no longer uses the Dae Jaweon, nor the Dai Kan-Wa Jiten, in its work. The Kangxi Dictionary and Hanyu Da Zidian are still used both in existing character source references, and as potential replacements for existing source references discovered to be erroneous. Similarly, although a (real or virtual) Kangxi Dictionary index was previously provided as part of the submission data for UTC-source characters, this is no longer the case. Instead, the stroke type of the first residual stroke (first stroke which does not form part of the radical) is supplied with all submitted characters, and used to order characters with the same radical and stroke count within the new Unicode block.

CJKV character 次 in traditional and simplified Chinese, Korean, Vietnamese and Japanese forms
CJKV character 次 in traditional and simplified Chinese, Korean, Vietnamese and Japanese forms

02CJK Unified Ideographs blocks

CJK Unified Ideographs

The basic block named CJK Unified Ideographs (4E00-9FFF) contains 20,992 basic Chinese characters in the range U+4E00 through U+9FFF. The block not only includes characters used in the Chinese writing system but also kanji used in the Japanese writing system, hanja in Korea, and chữ Nôm characters in Vietnamese. Many characters in this block are used in all three writing systems, while others are in only one or two of the three.

This block is also known as the Unified Repertoire and Ordering (URO), especially when it needs to be differentiated from the other CJK Unified Ideographs blocks.

The first 20,902 characters in the block are arranged according to the Kangxi Dictionary ordering of radicals. In this system the characters written with the fewest strokes are listed first. The remaining characters were added later, and so are not in radical order.

The block is the result of Han unification, which was somewhat controversial within East Asia. Since single characters used in more than one of Chinese, Japanese and Korean were coded in the same location, and the modern typographical conventions and handwriting curricula differ slightly between regions (not necessarily along language boundaries, for example, Hong Kong and Taiwan, which both use Traditional Chinese, have slightly different local conventions), the appearance of a selected glyph could depend on the particular font being used. However, the URO applies the source separation rule, meaning that pairs of characters treated as distinct in a character set used as a source for the URO (e.g. JIS X 0208 as used in e.g. Shift JIS) would remain pairs of separate characters in the new Unicode encoding.

Using variation selectors, it is possible to specify certain variant CJK ideograms within Unicode. The Adobe-Japan1 character set, which has 14,684 ideographic variation sequences, is an extreme example of the use of variation selectors.

Charts

4E00-62FF, 6300-77FF, 7800-8CFF, 8D00-9FFF.

Sources

Note: Most characters appear in multiple sources, so the sum of individual character counts (108,493) is far greater than the number of encoded characters (20,992).

MemberCodeSourceCharacter countTotal
ChinaG0GB/T 2312-1980 (formerly GB 2312-80)6,76320,938
G1GB/T 12345-1990 (formerly GB/T 12345-90); Traditional Chinese analogue to GB 2312-802,202
G3GB/T 13131 (unpublished GB/T 7589-1987 unsimplified forms)4,833
G5GB/T 13132 (unpublished GB/T 7590-1987 unsimplified forms)2,843
G7General Purpose Hanzi List for Modern Chinese Language, and General List of Simplified Hanzi42
G8GB/T 8565.2-1988 (formerly GB 8565.2-88)203
GCACulture and Art Publishing House Ideographs (文化艺术出版社用字)6
GCENames of newly-discovered chemical elements as assigned by the China National Committee for Terms in Sciences and Technologies
and the China National Language and Character Working Committee (全国科学技术名词审定委员会,国家语言文字工作委员会)
4
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China2
GEGB/T 16500-19983,767
GFCModern Chinese Standard Dictionary (现代汉语规范词典第二版)2
GGFZTongyong Guifan Hanzi Zidian (通用规范汉字字典)6
GGTCharacters collected by the National Library of China1
GHGB/T 15564-199559
GHZHanyu Da Zidian1
GHZRHanyu Da Zidian 2nd ed. (汉语大字典, 第二版)29
GKGB/T 12052-1989 (formerly GB 12052-89)89
GKJTerms in Sciences and Technologies approved by the China National Committee for Terms in Sciences and Technologies (CNCTST)16
GKXKangxi Dictionary43
GLKLongkan Shoujian1
GTStandard Telegraph Codebook (revised), 198316
GUNo source (the original source reference has been moved)8
GWYCultural Heritage Ideographs (文化遗产用字)1
GZFYHanyu Fangyan Dacidian (汉语方言大词典)1
Hong KongHHong Kong Supplementary Character Set, 20082,29215,376
HB0Computer Chinese Glyph and Character Code Mapping Table, Technical Report C-26
(電腦用中文字型與字碼對照表, 技術通報C-26)
9
HB1Big-5, Level 15,401
HB2Big-5, Level 27,650
HDHong Kong Supplementary Character Set, 201624
JapanJ0JIS X 0208-19906,35618,249
J1JIS X 0212-19903,058
J13JIS X 0213:2004 level-3 characters replacing J1 characters1,037
J13AJIS X 0213:2004 level-3 character addendum from JIS X 0213:2000 level-3 replacing J1 character2
J14JIS X 0213:2004 level-4 characters replacing J1 characters1,704
J3JIS X 0213:2004 Level 395
J3AJIS X 0213:2004 Level 3 addendum from JIS X 0213:2000 Level 37
J4JIS X 0213:2004 Level 4301
JARIBARIB STD-B24 Version 5.1, March 14 20073
JMJCharacter Information Development and Maintenance Project for e-Government "MojiJoho-Kiban Project" (文字情報基盤整備事業)5,686
North KoreaKP0KPS 9566-974,65215,008
KP1KPS 10721-200010,356
South KoreaK0KS X 1001:2004 (formerly KS C 5601-1987)4,62015,450
K1KS X 1002:2001 (formerly KS C 5657-1991)2,855
K2KS X 1027-1:2011 (formerly PKS C 5700-1 1994)7,911
K3KS X 1027-2:2011 (formerly PKS C 5700-2 1994)1
K4KS X 1027-3:2011 (formerly PKS 5700-3:1998)4
K6KS X 1027-5:202157
KCKorean History On-Line (한국 역사 정보 통합 시스템)1
KUNo source (the original source reference has been moved)1
MacauMAHKSCS-200829200
MB1Big Five10
MB2Big Five7
MCMacau Supplementary Character Set (MSCS) reference3
MDMacau Supplementary Character Set (MSCS) horizontal extensions127
MDHHKSCS-201624
TaiwanT1CNS 11643-1986 plane 15,41318,385
T2CNS 11643-1986 plane 27,651
T3CNS 11643-1992 plane 34,145
T4CNS 11643-1992 plane 4893
T5CNS 11643-1992 plane 564
T6CNS 11643-1992 plane 631
T7CNS 11643-1992 plane 716
TBCNS 11643-2007 plane 112
TCCNS 11643-2007 plane 122
TECNS 11643-2007 plane 149
TFCNS 11643-2007 plane 15159
VietnamV0TCVN 5773:19935984,808
V1TCVN 6056:19953,305
V2VHN 01-1998759
V3VHN 02-199891
V4Hán Nôm Coded Character Repertoire (Kho Chữ Hán Nôm Mã Hoá)19
VNVietnamese horizontal and vertical extensions36
N/AUTCUTC sources7979

In Unicode 4.1, 14 HKSCS-2004 characters and 8 GB 18030 characters were assigned to between U+9FA6 and U+9FBB code points. Since then, other additions were added to this block for various reasons, all summarized in the version history section below.

CJK Unified Ideographs Extension A

The block named CJK Unified Ideographs Extension A (3400-4DBF) contains 6,592 additional characters in the range U+3400 through U+4DBF.

Charts

3400-4DBF.

Sources

Note: Most characters appear in more than one source, so the sum of individual character counts (23,997) is far greater than the number of encoded characters (6,592).

MemberCodeSourceCharacter countTotal
 ChinaG3GB/T 13131 (unpublished GB/T 7589-1987 unsimplified forms)2,3906,230
G5GB/T 13132 (unpublished GB/T 7590-1987 unsimplified forms)1,226
G7General Purpose Hanzi List for Modern Chinese Language, and General List of Simplified Hanzi120
GCACulture and Art Publishing House Ideographs12
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China7
GGFZTongyong Guifan Hanzi Zidian2
GHZHanyu Da Zidian341
GHZRHanyu Da Zidian 2nd ed.1
GKJTerms in Sciences and Technologies approved by the China National Committee for Terms in Sciences and Technologies (CNCTST)3
GKXKangxi Dictionary1,889
GSSingapore Chinese characters226
GWYCultural Heritage Ideographs1
GZAncient Zhuang Character Dictionary12
 Hong KongHHong Kong Supplementary Character Set, 2008572572
 JapanJ3JIS X 0213:2004 Level 325,856
J4JIS X 0213:2004 Level 478
JAJapanese IT Vendors Contemporary Ideographs, 1993574
JA3JIS X 0213:2004 level-3 characters replacing JA characters17
JA4JIS X 0213:2004 level-4 characters replacing JA characters67
JMJCharacter Information Development and Maintenance Project for e-Government "MojiJoho-Kiban Project"5,118
 North KoreaKP0KPS 9566-9713,191
KP1KPS 10721-20003,190
 South KoreaK3KS X 1027-2:2011 (formerly PKS C 5700-2 1994)1,8331,876
K4KS X 1027-3:2011 (formerly PKS 5700-3:1998)2
K6KS X 1027-5:202137
KCKorean History On-Line3
KUNo source (the original source reference has been moved)1
 MacauMAHKSCS-2008412
MDMacau Supplementary Character Set (MSCS) horizontal extensions8
 TaiwanT3CNS 11643-1992 plane 32,1795,917
T4CNS 11643-1992 plane 42,920
T5CNS 11643-1992 plane 5400
T6CNS 11643-1992 plane 6200
T7CNS 11643-1992 plane 7133
TECNS 11643-2007 plane 141
TFCNS 11643-2007 plane 1584
United KingdomUKIRG N2107R233
 VietnamV0TCVN 5773:1993140319
V2VHN 01-1998149
V3VHN 02-199819
V4Hán Nôm Coded Character Repertoire5
VNVietnamese horizontal and vertical extensions6
N/AUTCUTC sources2121

CJK Unified Ideographs Extension B

The block named CJK Unified Ideographs Extension B (20000-2A6DF) contains 42,720 characters in the range U+20000 through U+2A6DF. These include most of the characters used in the Kangxi Dictionary that are not in the basic CJK Unified Ideographs block, as well as many Hán-Nôm characters that were formerly used to write Vietnamese.

Charts

20000-215FF, 21600-230FF, 23100-245FF, 24600-260FF, 26100-275FF, 27600-290FF, 29100-2A6DF.

Sources

Note: Many characters appear in more than one source, so the sum of individual character counts (100,887) is far greater than the number of encoded characters (42,720).

MemberCodeSourceCharacter countTotal
 ChinaG3GB/T 13131 (unpublished GB/T 7589-1987 unsimplified forms)131,345
G4KSiku Quanshu474
GBKEncyclopedia of China59
GCACulture and Art Publishing House Ideographs78
GCESICharacters collected by China Electronics Standardization Institute (中国电子技术标准化研究院)102
GCHCihai247
GCYCiyuan66
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China146
GFZFounder Press System (方正排版系统)65
GGFZTongyong Guifan Hanzi Zidian5
GHCHanyu Da Cidian553
GHFHanwen fodian yinan suzi huishi yu yanjiu (漢文佛典疑難俗字彙釋與研究)1
GHZHanyu Da Zidian10,506
GHZRHanyu Da Zidian 2nd ed.4
GKJTerms in Sciences and Technologies approved by the China National Committee for Terms in Sciences and Technologies (CNCTST)17
GKXKangxi Dictionary18,472
GUNo source (the original source reference has been moved)73
GWYCultural Heritage Ideographs12
GXMCharacters for use in personal names in China from Public Order Administration, Ministry of Public Security of the People's Republic of China8
GZAncient Zhuang Character Dictionary (古壮字字典)453
GZFYHanyu Fangyan Dacidian3
 Hong KongHHong Kong Supplementary Character Set, 20081,7031,703
 JapanJ3JIS X 0213:2004 Level 32525,745
J3AJIS X 0213:2004 Level 3 addendum from JIS X 0213:2000 Level 31
J4JIS X 0213:2004 Level 4277
JMJCharacter Information Development and Maintenance Project for e-Government "MojiJoho-Kiban Project"25,442
 North KoreaKP1KPS 10721-20005,7665,766
 South KoreaK1KS X 1002:2001 (formerly KS C 5657-1991)1683
K4KS X 1027-3:2011 (formerly PKS 5700-3:1998)166
K6KS X 1027-5:2021502
KCKorean History On-Line14
 MacauMAHKSCS-2008938
MCMacau Supplementary Character Set (MSCS) reference2
MDMacau Supplementary Character Set (MSCS) horizontal extensions27
 TaiwanT3CNS 11643-1992 plane 32830,212
T4CNS 11643-1992 plane 43,408
T5CNS 11643-1992 plane 58,114
T6CNS 11643-1992 plane 65,942
T7CNS 11643-1992 plane 76,299
TACNS 11643-2007 plane 1012
TBCNS 11643-2007 plane 117
TCCNS 11643-2007 plane 121
TFCNS 11643-2007 plane 156,401
 United KingdomUKIRG N2107R21212
 VietnamV0TCVN 5773:19931,5705,299
V1TCVN 6056:19951
V2VHN 01-19982,286
V3VHN 02-1998422
V4Hán Nôm Coded Character Repertoire33
VNVietnamese horizontal and vertical extensions987
Buddhist canonSATSAT Daizōkyō Text Database11
N/AUTCUTC sources8383

CJK Unified Ideographs Extension C

The block named CJK Unified Ideographs Extension C (2A700-2B73F) contains 4,160 characters in the range U+2A700 through U+2B73F. It was initially added in Unicode 5.2 (2009).

Charts

2A700-2B73F.

Sources

Note: Some characters appear in more than one source, so the sum of individual character counts (4,967) is greater than the number of encoded characters (4,160).

MemberCodeSourceCharacter countTotal
 ChinaGBKEncyclopedia of China741,456
GCACulture and Art Publishing House Ideographs12
GCESICharacters collected by China Electronics Standardization Institute117
GCHCihai264
GCYCiyuan1
GCYYChinese Academy of Surveying and Mapping ideographs (中国测绘科学院用字)55
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China83
GFZFounder Press System1
GGFZTongyong Guifan Hanzi Zidian2
GGHGudai Hanyu Cidian (古代汉语词典)51
GHCHanyu Da Cidian14
GHZHanyu Da Zidian1
GHZRHanyu Da Zidian 2nd ed.1
GJZCommercial Press ideographs (商务印书馆用字)61
GKJTerms in Sciences and Technologies approved by the China National Committee for Terms in Sciences and Technologies (CNCTST)6
GKXKangxi Dictionary8
GWYCultural Heritage Ideographs2
GXCXiandai Hanyu Cidian25
GXMCharacters for use in personal names in China from Public Order Administration, Ministry of Public Security of the People's Republic of China2
GZAncient Zhuang Character Dictionary109
GZFYHanyu Fangyan Dacidian202
GZJWYin Zhou Jinwen Jicheng Yinde (殷周金文集成引得)365
 Hong KongHHong Kong Supplementary Character Set, 200811
 JapanJKJapanese Kokuji Collection (Mojikyō subset)367431
JMJCharacter Information Development and Maintenance Project for e-Government "MojiJoho-Kiban Project"64
 North KoreaKP1KPS 10721-200099
 South KoreaK5KS X 1027-4:2011 (formerly Korean IRG Hanja Character Set 5th Edition: 2001)404407
K6KS X 1027-5:20212
KCKorean History On-Line1
 MacauMCMacau Supplementary Character Set (MSCS) reference1721
MDMacau Supplementary Character Set (MSCS) horizontal extensions4
 TaiwanT4CNS 11643-1992 plane 411,757
T5CNS 11643-1992 plane 51
T6CNS 11643-1992 plane 62
TBCNS 11643-2007 plane 112
TCCNS 11643-2007 plane 12634
TDCNS 11643-2007 plane 13766
TECNS 11643-2007 plane 14350
TUNo source (the original source reference has been moved)1
 United KingdomUKIRG N2107R211
 VietnamV0TCVN 5773:19934795
V1TCVN 6056:19952
V2VHN 01-19981
V4Hán Nôm Coded Character Repertoire782
VNVietnamese horizontal and vertical extensions6
N/AUTCUTC sources8989

CJK Unified Ideographs Extension D

The block named CJK Unified Ideographs Extension D (2B740-2B81F) contains 222 characters in the range U+2B740 through U+2B81D that were added in Unicode 6.0 (2010).

Charts

2B740-2B81F.

Sources

Note: Some characters appear in more than one source, so the sum of individual character counts (260) is greater than the number of encoded characters (222).

MemberCodeSourceCharacter countTotal
 ChinaGCACulture and Art Publishing House Ideographs1299
GCESICharacters collected by China Electronics Standardization Institute6
GCHCihai1
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China1
GIDCID System of the Ministry of Public Security of China (公安人口信息专用字库补充汉字)9
GKJTerms in Sciences and Technologies approved by the China National Committee for Terms in Sciences and Technologies (CNCTST)2
GXCXiandai Hanyu Cidian4
GXMCharacters for use in personal names in China from Public Order Administration, Ministry of Public Security of the People's Republic of China22
GZAncient Zhuang Character Dictionary3
GZHZhonghua Zihai39
 JapanJHHanyo-Denshi Program (汎用電子情報交換環境整備プログラム)107117
JMJCharacter Information Development and Maintenance Project for e-Government "MojiJoho-Kiban Project"10
 TaiwanTBCNS 11643-2007 plane 112424
N/AUTCUTC sources2020

CJK Unified Ideographs Extension E

The block named CJK Unified Ideographs Extension E (2B820-2CEAF) contains 5,774 characters in the range U+2B820 through U+2CEAD. It was originally added in Unicode 8.0 (2015).

Charts

2B820-2CEAF.

Sources

Note: Some characters appear in more than one source, so the sum of individual character counts (6,272) is greater than the number of encoded characters (5,774).

MemberCodeSourceCharacter countTotal
 ChinaGBKEncyclopedia of China113,173
GCACulture and Art Publishing House Ideographs20
GCESICharacters collected by China Electronics Standardization Institute211
GCHCihai112
GCYCiyuan3
GCYYChinese Academy of Surveying and Mapping ideographs98
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China10
GDZGeographic Publishing House ideographs (地质出版社用字)1
GGFZTongyong Guifan Hanzi Zidian4
GGHGudai Hanyu Cidian175
GGTCharacters collected by the National Library of China2
GHCHanyu Da Cidian (漢語大詞典)7
GIDCID System of the Ministry of Public Security of China37
GJZCommercial Press ideographs147
GKJTerms in Sciences and Technologies approved by the China National Committee for Terms in Sciences and Technologies (CNCTST)2
GKXKangxi Dictionary22
GRMPeople's Daily ideographs (人民日报用字)3
GUNo source (the original source reference has been moved)3
GWYCultural Heritage Ideographs1
GWZHanyu Da Cidian Press ideographs (漢語大詞典出版社用字)12
GXCXiandai Hanyu Cidian57
GXHXinhua Dictionary4
GXMCharacters for use in personal names in China from Public Order Administration, Ministry of Public Security of the People's Republic of China1
GZAncient Zhuang Character Dictionary107
GZFYHanyu Fangyan Dacidian712
GZHSJCharacters collected by the Zhonghua Book Company (中华书局)1
GZJWYin Zhou Jinwen Jicheng Yinde1,410
 Hong KongHDHong Kong Supplementary Character Set, 201611
 JapanJKJapanese Kokuji Collection (Mojikyō subset)415503
JMJCharacter Information Development and Maintenance Project for e-Government "MojiJoho-Kiban Project"88
 South KoreaKCKorean History On-Line77
 MacauMCMacau Supplementary Character Set (MSCS) reference4851
MDMacau Supplementary Character Set (MSCS) horizontal extensions3
 TaiwanT3CNS 11643-1992 plane 321,261
TBCNS 11643-2007 plane 112
TCCNS 11643-2007 plane 12323
TDCNS 11643-2007 plane 13595
TECNS 11643-2007 plane 14339
 United KingdomUKIRG N2107R222
 VietnamV0TCVN 5773:199371,037
V2VHN 01-19981
V4Hán Nôm Coded Character Repertoire1,023
VNVietnamese horizontal and vertical extensions6
N/AUTCUTC sources237237

CJK Unified Ideographs Extension F

The block named CJK Unified Ideographs Extension F (2CEB0-2EBEF) contains 7,473 characters in the range U+2CEB0 through 2EBE0 that were added in Unicode 10.0 (2017). It includes more than 1,000 Sawndip characters for Zhuang.

Charts

2CEB0-2EBEF.

Sources

Note: Some characters appear in more than one source, so the sum of individual character counts (8,015) is greater than the number of encoded characters (7,473).

MemberCodeSourceCharacter countTotal
 ChinaGCACulture and Art Publishing House Ideographs461,546
GCESICharacters collected by China Electronics Standardization Institute73
GCYCiyuan122
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China31
GFCModern Chinese Standard Dictionary27
GIDCID System of the Ministry of Public Security of China1
GKJTerms in Sciences and Technologies approved by the China National Committee for Terms in Sciences and Technologies (CNCTST)5
GLGYJZhuang Liao Songs Research (壮族嘹歌研究)1
GOCDOxford English-Chinese Chinese-English Dictionary (牛津英汉汉英词典)2
GPGLGZhuang Folk Song Culture Series - Pingguo County Liao Songs (壮族民歌文化丛书•平果嘹歌)69
GWYCultural Heritage Ideographs6
GXHZXinhua Da Zidian (新华大字典)51
GXMCharacters for use in personal names in China from Public Order Administration, Ministry of Public Security of the People's Republic of China2
GZAncient Zhuang Character Dictionary1,075
GZJWYin Zhou Jinwen Jicheng Yinde33
GZYSChinese Ancient Ethnic Characters Research, 1984 (中国民族古文字研究)2
 Hong KongHDHong Kong Supplementary Character Set, 201611
 JapanJMJCharacter Information Development and Maintenance Project for e-Government "MojiJoho-Kiban Project"1,6461,646
 South KoreaKCKorean History On-Line1,8101,810
 MacauMCMacau Supplementary Character Set (MSCS) reference2222
 TaiwanT3CNS 11643-1992 plane 316
T6CNS 11643-1992 plane 62
T7CNS 11643-1992 plane 72
TCCNS 11643-2007 plane 121
 United KingdomUKIRG N2107R222
 VietnamV0TCVN 5773:1993117
V4Hán Nôm Coded Character Repertoire8
VNVietnamese horizontal and vertical extensions8
Buddhist canonSATSAT Daizōkyō Text Database2,8842,884
N/AUTCUTC sources8181

CJK Unified Ideographs Extension G

A block named CJK Unified Ideographs Extension G was added as part of Unicode 13.0 to the Tertiary Ideographic Plane in the range U+30000 through U+3134F, containing 4,939 characters.

Charts

30000-3134F.

Sources

Note: Some characters appear in more than one source, so the sum of individual character counts (5,239) is greater than the number of encoded characters (4,939).

MemberCodeSourceCharacter countTotal
 ChinaGCACulture and Art Publishing House Ideographs692,239
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China49
GHZRHanyu Da Zidian 2nd ed.878
GPGLGZhuang Folk Song Culture Series - Pingguo County Liao Songs13
GWYCultural Heritage Ideographs11
GXMCharacters for use in personal names in China from Public Order Administration, Ministry of Public Security of the People's Republic of China11
GZAncient Zhuang Character Dictionary1,208
 South KoreaKCKorean History On-Line435435
 TaiwanT13CNS 11643 (pending new version) plane 19347354
T5CNS 11643-1992 plane 51
TBCNS 11643-2007 plane 113
TCCNS 11643-2007 plane 122
TDCNS 11643-2007 plane 131
 United KingdomUKIRG N2107R21,5661,566
 VietnamV4Hán Nôm Coded Character Repertoire676
VNVietnamese horizontal and vertical extensions70
Buddhist canonSATSAT Daizōkyō Text Database329329
N/AUTCUTC sources240240

CJK Unified Ideographs Extension H

A block named CJK Unified Ideographs Extension H was added as part of Unicode 15.0 to the Tertiary Ideographic Plane in the range U+31350 through U+323AF, containing 4,192 characters.

Charts

31350-323AF.

Sources

Note: Some characters appear in more than one source, so the sum of individual character counts (4,541) is greater than the number of encoded characters (4,192).

MemberCodeSourceCharacter countTotal
 ChinaGCACulture and Art Publishing House Ideographs91,059
GCESICharacters collected by China Electronics Standardization Institute1
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China298
GHCHanyu Da Cidian27
GKJTerms in Sciences and Technologies approved by the China National Committee for Terms in Sciences and Technologies (CNCTST)30
GLGYJZhuang Liao Songs Research (壮族嘹歌研究)11
GPGLGZhuang Folk Song Culture Series - Pingguo County Liao Songs (壮族民歌文化丛书•平果嘹歌)14
GUNo source (the original source reference has been moved)1
GWYCultural Heritage Ideographs5
GXMCharacters for use in personal names in China from Public Order Administration, Ministry of Public Security of the People's Republic of China216
GZAncient Zhuang Character Dictionary330
GZA-1A Vibrant and Unbroken Transmission, Filial Piety and Zhuang Funeral Songs (生生不息的传承•孝与壮族行孝歌之研究)6
GZA-2Annotated Long Zhuang Morality Songs (壮族伦理道德长诗传扬歌译注)38
GZA-3Compendium of Old Zhuang Folksong Texts, Wooing Songs vol. 1, Liao Songs (壮族民歌古籍集成•情歌(一)嘹歌)2
GZA-4Compendium of Old Zhuang Folksong Texts, Wooing Songs vol. 2, Fwen Nganx (壮族民歌古籍集成•情歌(二)欢𭪤)11
GZA-6Zhuang Proverbs from China (中国壮族谚语)59
GZA-7Ancient Remembrance, Zhuang Creation Myth Songs (远古的追忆•壮族创世神话古歌研究)1
 North KoreaKP1KPS 10721-200011
 South KoreaKCKorean History On-Line512512
 TaiwanT12CNS 11643 (pending new version) plane 187716
T13CNS 11643 (pending new version) plane 19696
T4CNS 11643-1992 plane 41
T6CNS 11643-1992 plane 61
T7CNS 11643-1992 plane 72
TBCNS 11643-2007 plane 115
TCCNS 11643-2007 plane 123
TECNS 11643-2007 plane 141
 United KingdomUKIRG N2232R917917
 VietnamV0TCVN 5773:19936931
V4Hán Nôm Coded Character Repertoire74
VNVietnamese horizontal and vertical extensions851
Buddhist canonSATSAT Daizōkyō Text Database241241
N/AUTCUTC sources164164

CJK Unified Ideographs Extension I

A block named CJK Unified Ideographs Extension I was added as part of Unicode 15.1 to the Supplementary Ideographic Plane in the range U+2EBF0 through U+2EE5F, containing 622 characters.

Charts

2EBF0-2EE5F.

Sources

Note: Some characters appear in more than one source, making the sum of individual character counts (625) more than the number of encoded characters (622).

MemberCodeSourceCharacter countTotal
 ChinaGIDC23ID system of the Ministry of Public Security of China, 2023622622
 JapanJMJCharacter Information Development and Maintenance Project for e-Government "MojiJoho-Kiban Project"11
N/AUTCUTC sources22

CJK Unified Ideographs Extension J

A block named CJK Unified Ideographs Extension J was added as part of Unicode 17.0 to the Tertiary Ideographic Plane in the range U+323B0-U+33479, containing 4,298 characters.

Charts

323B0-3347F.

Sources

Note: Some characters appear in more than one source, making the sum of individual character counts (4,406) more than the number of encoded characters (4,298).

MemberCodeSourceCharacter countTotal
 ChinaGDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China1441,005
GKJTerms in Sciences and Technologies approved by the China National Committee for Terms in Sciences and Technologies (CNCTST)567
GXMCharacters for use in personal names in China from Public Order Administration, Ministry of Public Security of the People's Republic of China4
GZAncient Zhuang Character Dictionary290
 South KoreaKCKorean History On-Line178178
 TaiwanT11CNS 11643 (pending new version) plane 171937
T9CNS 11643 (pending new version) plane 959
TBCNS 11643-2007 plane 1169
TCCNS 11643-2007 plane 12165
TDCNS 11643-2007 plane 13241
TECNS 11643-2007 plane 14396
TFCNS 11643-2007 plane 156
 United KingdomUKIRG N2232R6906
UKIRG N2487900
 VietnamV0TCVN 5773:199316991
V1TCVN 6056:19951
V2VHN 01-19982
V3VHN 02-19981
V4Hán Nôm Coded Character Repertoire58
VNVietnamese horizontal and vertical extensions913
Buddhist canonSATSAT Daizōkyō Text Database259260
SATMSAT manuscript collection for Buddhist studies1
N/AUTCUTC sources129129

CJK Compatibility Ideographs

The block named CJK Compatibility Ideographs (F900-FAFF) was created to retain round-trip compatibility with other standards.

However, twelve characters in this block actually have the "Unified Ideograph" property: U+FA0E 﨎, U+FA0F 﨏, U+FA11 﨑, U+FA13 﨓, U+FA14 﨔, U+FA1F 﨟, U+FA21 﨡, U+FA23 﨣, U+FA24 﨤, U+FA27 﨧, U+FA28 﨨, and U+FA29 﨩. None of the other characters in this and other "Compatibility" blocks relate to CJK unification.

While 龜 and 亀 are not considered unifiable, U+FA20 CJK COMPATIBILITY IDEOGRAPH-FA20 is considered a duplicate to U+8612 CJK UNIFIED IDEOGRAPH-8612.

Charts

F900-FAFF.

Sources

Note: All characters appear in more than one source, so the sum of individual character counts () is greater than the number of encoded characters ().

MemberCodeSourceCharacter countTotal
 ChinaGCACulture and Art Publishing House Ideographs425
GDMPlace name characters from the Public Order Administration, Ministry of Public Security of the People's Republic of China1
GUNo source (the original source reference has been moved)20
 Hong KongHHong Kong Supplementary Character Set, 200813
HB2Big-5, Level 22
 JapanJ3JIS X 0213:2004 Level 372101
J4JIS X 0213:2004 Level 49
JAJapanese IT Vendors Contemporary Ideographs, 19931
JA3JIS X 0213:2004 level-3 characters replacing JA characters1
JARIBARIB STD-B24 Version 5.1, March 14 20073
JMJCharacter Information Development and Maintenance Project for e-Government "MojiJoho-Kiban Project"15
 North KoreaKP1KPS 10721-2000106107
KPUThe source reference for this ideograph has been moved; the value is its code point.1
 South KoreaK0KS X 1001:2004 (formerly KS C 5601-1987)270270
 TaiwanTFCNS 11643-2007 plane 1511
 VietnamV0TCVN 5773:199333
N/AUTCUTC sources3434

03Known issues

Disunification

U+4039

The character U+4039 (䀹) was a unification of two different characters (one with jiā 夾 phonetic and one with shǎn 㚒 phonetic) until Unicode 5.0. However, they were lexically different characters that should not have been unified; they have different pronunciations and different meanings.

The proposal of disunification of U+4039 was accepted for Unicode 5.1, encoding a new character at U+9FC3 (鿃) to represent shǎn.

Other 3 glyphs in Extension B

In CJK Unified Ideographs Extension B, some characters were incorrectly unified with others. These characters include U+2017B (𠅻), U+204AF (𠒯) and U+24CB2 (𤲲). The first two characters contained a wrong unification of Chinese and Vietnamese source of their glyph, while the last one unifies the Chinese and Taiwanese ones.

The glyphs for U+2017B (𠅻) and U+204AF (𠒯) were corrected in version 10.0, and the erroneous UCS2003 source glyph U+24CB2 (𤲲) was removed in version 13.0.

Unifiable variants and exact duplicates

Also in CJK Unified Ideographs Extension B, hundreds of glyph variants were encoded by mistake. Additionally, an ISO/IEC JTC 1/SC 2 report has found that six exact duplicates (where the same character has inadvertently been encoded twice) and two semi-duplicates (where the CJK-B character represents a de facto disunification of two glyph forms unified in the corresponding BMP character) were encoded by mistake:

  • U+34A8 㒨 = U+20457 𠑗 : U+20457 is the same as the China-source glyph for U+34A8, but it is significantly different from the Taiwan-source glyph for U+34A8
  • U+3DB7 㶷 = U+2420E 𤈎 : same glyph shapes
  • U+8641 虁 = U+27144 𧅄 : U+27144 is the same as the Korean-source glyph for U+8641, but it is significantly different from the mainland China-, Taiwan- and Japan-source glyphs for U+8641
  • U+204F2 𠓲 = U+23515 𣔕 : same glyph shapes, but ordered under different radicals
  • U+249BC 𤦼 = U+249E9 𤧩 : same glyph shapes
  • U+24BD2 𤯒 = U+2A415 𪐕 : same glyph shapes, but ordered under different radicals
  • U+26842 𦡂 = U+26866 𦡦 : same glyph shapes
  • U+FA23 﨣 = U+27EAF 𧺯 : same glyph shapes (U+FA23 﨣 is a unified CJK ideograph, despite its name "CJK COMPATIBILITY IDEOGRAPH-FA23.")

04Other CJK ideographs in Unicode, not Unified

Apart from the eleven blocks of "Unified Ideographs," Unicode has about a dozen more blocks with not-unified CJK-characters. These are mainly CJK radicals, strokes, punctuation, marks, symbols and compatibility characters. Although some characters have their (decomposable) counterparts in other blocks, the usages can be different. An example of a not-unified CJK-character is U+3007 IDEOGRAPHIC NUMBER ZERO in the CJK Symbols and Punctuation block. Although it is not covered under "CJK Unified Ideographs", it is treated as a CJK-character for all other intents and purposes.

Four blocks of compatibility characters are included for compatibility with legacy text handling systems and older character sets:

They include forms of characters for vertical text layout and rich text characters that Unicode recommends handling through other means. Therefore, their use is discouraged.

05Unihan database

The Unihan database maintained by the Unicode Consortium provides information about all of the unified Han characters encoded in the Unicode Standard, including mappings to various national and industry standards, indices into standard dictionaries, encoded variants, pronunciations in various languages, and an English definition. The database is available to the public as text files and via an interactive website. The latter also includes representative glyphs and definitions for compound words drawn from the free Japanese EDICT and Chinese CEDICT dictionary projects (which are provided for convenience and are not a formal part of the Unicode Standard).

06Font support

The blocks CJK Unified Ideographs and CJK Unified Ideographs Extension A, being parts of the Basic Multilingual Plane, are supported by the majority of the CJK fonts. However, Japanese and Korean fonts usually have fewer characters (about 13,000 and 8,000, respectively) than Chinese. Extensions B, C, D are supported by additional fonts MingLiU-ExtB, MingLiU_HKSCS-ExtB, PMingLiU-ExtB, SimSun-ExtB included in Microsoft Windows since Vista.

07Unicode version history

CJK unified ideograph additions per Unicode version
Unicode versionAdditionPlaneCharacters addedTotal characters
1.0 (1991)CJK Compatibility IdeographsBasic Multilingual Plane (BMP)1220,914
CJK Unified IdeographsBMP20,902
3.0 (1999)CJK Unified Ideographs Extension ABMP6,58227,496
3.1 (2001)CJK Unified Ideographs Extension BSupplementary Ideographic Plane (SIP)42,71170,207
4.1 (2005)CJK Unified Ideographs: Ideographs from HKSCS-2004 and GB 18030-2000 not in ISO 10646BMP2270,229
5.1 (2008)CJK Unified Ideographs: Ideographs from Adobe Japan and disunification of U+4039BMP870,237
5.2 (2009)CJK Unified Ideographs: Characters from ARIB #47, #95, #93 and HKSCSBMP874,394
CJK Unified Ideographs Extension CSIP4,149
6.0 (2010)CJK Unified Ideographs Extension DSIP22274,616
6.1 (2012)CJK Unified Ideographs: Character corresponding to Adobe-Japan1-6 CID+20156BMP174,617
8.0 (2015)CJK Unified IdeographsBMP980,388
CJK Unified Ideographs Extension ESIP5,762
10.0 (2017)CJK Unified IdeographsBMP2187,882
CJK Unified Ideographs Extension FSIP7,473
11.0 (2018)CJK Unified IdeographsBMP587,887
13.0 (2020)CJK Unified IdeographsBMP1392,856
CJK Unified Ideographs Extension ABMP10
CJK Unified Ideographs Extension BSIP7
CJK Unified Ideographs Extension GTertiary Ideographic Plane (TIP)4,939
14.0 (2021)CJK Unified IdeographsBMP392,865
CJK Unified Ideographs Extension BSIP2
CJK Unified Ideographs Extension CSIP4
15.0 (2022)CJK Unified Ideographs Extension CSIP197,058
CJK Unified Ideographs Extension HTIP4,192
15.1 (2023)CJK Unified Ideographs Extension ISIP62297,680
17.0 (2025)CJK Unified Ideographs Extension CSIP6101,996
CJK Unified Ideographs Extension ESIP12
CJK Unified Ideographs Extension JTIP4,298
Watch videos about CJK Unified IdeographsExplainers and documentaries on YouTube (opens in a new tab)

Sources and credits

This article is adapted from the Wikipedia article CJK Unified Ideographs, written by its contributors and licensed under CC BY-SA 4.0. Fathomly has changed the layout, removed citation markers, navigation and maintenance notices, and adjusted punctuation. This adapted version is shared under the same license. For references, see the original article.

Images, from Wikimedia Commons:

Fathomly is not affiliated with or endorsed by the Wikimedia Foundation. Spotted a problem? Tell us.