[PATCH 1/3] contrib, libcpp, libstdc++: Update to Unicode 18.0.0
Jonathan Wakely
jwakely@redhat.com
Thu Sep 17 10:56:31 GMT 2026
On Thu, 17 Sept 2026 at 11:04, Jakub Jelinek <jakub@redhat.com> wrote:
>
> Hi!
>
> Unicode 18.0.0 has been released yesterday.
> The following patch follows the unicode/README and updates
> whatever is needed in gcc to that Unicode version. In glibc there
> weren't any changes to the from_glibc files.
>
> Note, apparently I forgot last year to bump in libstdc++ the
> inline namespace and 160000 code to 170000, this bumps it right to
> 180000. As for tests, I've added some new characters to
> named-universal-char-escape-1.c and made sure to test boundaries of
> generated ranges where it changed and test new generated ranges.
> In libstdc++ testsuite/ext/unicode/properties.cc test I had to tweak
> one static assert because EastAsianWidth.txt changed for the character
> being tested (and there is no character right above U+1FAFA).
>
> As usually, the patch is way too large for the mailing list limits
> even when xz -9e compressed.
> Below is included the ChangeLog and the non-generated changes, attached
> is xz -9e compressed first part of the generated changes (all
> generated files except the huge uname2c.h second table), in a follow-up
> I'll post one xz -9e compressed patch which removes the uname2c.h
> second table and another one which adds it again with the new content
> (basically the whole array is changed, so a diff which replaces it
> is alone also too large).
>
> So far lightly tested, ok for trunk if it passes full bootstrap/regtest?
The libstdc++ parts are OK, thanks.
>
> 2026-09-17 Jakub Jelinek <jakub@redhat.com>
>
> contrib/
> * unicode/README: Replace Unicode 17 version with 18.
> * unicode/gen_libstdcxx_unicode_data.py: Use 180000 instead of
> 160000 in _GLIBCXX_GET_UNICODE_DATA test.
> * unicode/DerivedCoreProperties.txt: Updated from Unicode 18.0.
> * unicode/emoji-data.txt: Likewise.
> * unicode/PropList.txt: Likewise.
> * unicode/HangulSyllableType.txt: Likewise.
> * unicode/DerivedNormalizationProps.txt: Likewise.
> * unicode/NameAliases.txt: Likewise.
> * unicode/UnicodeData.txt: Likewise.
> * unicode/EastAsianWidth.txt: Likewise.
> gcc/testsuite/
> * c-c++-common/cpp/named-universal-char-escape-1.c: Add tests
> for some Unicode 18.0 characters, both normal and generated.
> libcpp/
> * makeucnid.cc (write_copyright): Update Unicode Copyright years.
> * makeuname2c.cc (generated_ranges): Adjust Unicode version from 17.0
> to 18.0. Add JURCHEN CHARACTER- and SMALL SEAL CHARACTER- generated
> ranges, adjust indexes in following entries.
> (write_copyright): Update Unicode Copyright years.
> * generated_cpp_wcwidth.h: Regenerated.
> * ucnid.h: Regenerated.
> * uname2c.h: Regenerated.
> libstdc++-v3/
> * include/bits/unicode.h (std::__unicode::__v16_0_0): Rename inline
> namespace to ...
> (std::__unicode::__v18_0_0): ... this.
> (_GLIBCXX_GET_UNICODE_DATA): Change from 160000 to 180000.
> * testsuite/ext/unicode/properties.cc: Expect 2 rather than 1 for
> uc::__field_width(U'\U0001FAF9').
> * include/bits/unicode-data.h: Regenerated.
>
> --- a/contrib/unicode/README 2026-08-06 11:38:29.988537277 +0200
> +++ b/contrib/unicode/README 2026-09-17 10:10:21.815296372 +0200
> @@ -47,14 +47,14 @@ produce ucnid.h.
>
> The procedure to update GCC's Unicode support is the following:
>
> -1. Update the six Unicode data files from the above URLs.
> +1. Update the ten Unicode data files from the above URLs.
>
> 2. Update the two glibc files in from_glibc/ from glibc's git. Update
> the commit number above in this README.
>
> 3. Run ./gen_wcwidth.py X.Y > ../../libcpp/generated_cpp_wcwidth.h
> (where X.Y is the version of the Unicode standard corresponding to the
> - Unicode data files being used, most recently, 17.0.0).
> + Unicode data files being used, most recently, 18.0.0).
>
> 4. Update Unicode Copyright years in libcpp/makeucnid.cc and in
> libcpp/makeuname2c.cc up to the year in which the Unicode
> @@ -69,7 +69,7 @@ The procedure to update GCC's Unicode su
> PropList.txt > ../../libcpp/ucnid.h
>
> 7. Read the corresponding Unicode's standard and update correspondingly
> - generated_ranges table in libcpp/makeuname2c.cc (in Unicode 17 all
> + generated_ranges table in libcpp/makeuname2c.cc (in Unicode 18 all
> the needed information was in Table 4-8).
>
> 8. Compile makeuname2c, e.g. with:
> --- a/contrib/unicode/gen_libstdcxx_unicode_data.py 2026-03-27 10:17:13.244345262 +0100
> +++ b/contrib/unicode/gen_libstdcxx_unicode_data.py 2026-09-17 10:29:44.627380770 +0200
> @@ -64,7 +64,7 @@ print("""
> """)
> print("#ifndef _GLIBCXX_GET_UNICODE_DATA")
> print('# error "This is not a public header, do not include it directly"')
> -print("#elif _GLIBCXX_GET_UNICODE_DATA != 160000")
> +print("#elif _GLIBCXX_GET_UNICODE_DATA != 180000")
> print('# error "Version mismatch for Unicode static data"')
> print("#endif\n")
>
> --- a/gcc/testsuite/c-c++-common/cpp/named-universal-char-escape-1.c 2026-03-27 10:17:15.174313768 +0100
> +++ b/gcc/testsuite/c-c++-common/cpp/named-universal-char-escape-1.c 2026-09-17 10:59:47.165738592 +0200
> @@ -119,6 +119,9 @@ typedef __CHAR32_TYPE__ char32_t;
> || U'\uFE18' != U'\N{PRESENTATION FORM FOR VERTICAL RIGHT WHITE LENTICULAR BRACKET}' \
> || U'\u0CF3' != U'\N{KANNADA SIGN COMBINING ANUSVARA ABOVE RIGHT}' \
> || U'\u0ECE' != U'\N{LAO YAMAKKAN}' \
> + || U'\u05C9' != U'\N{HEBREW POINT DAGESH HAZAQ MUDGASH}' \
> + || U'\u20C3' != U'\N{UAE DIRHAM SIGN}' \
> + || U'\U00010EE1' != U'\N{ARABIC CROWN LETTER SHEEN}' \
> || U'\U00010EFE' != U'\N{ARABIC SMALL LOW WORD QASR}' \
> || U'\U00011241' != U'\N{KHOJKI VOWEL SIGN VOCALIC R}' \
> || U'\U00011B06' != U'\N{DEVANAGARI SIGN WESTERN FIVE-LIKE BHALE}' \
> @@ -160,6 +163,7 @@ typedef __CHAR32_TYPE__ char32_t;
> || U'\U0002B739' != U'\N{CJK UNIFIED IDEOGRAPH-2B739}' \
> || U'\U0002B740' != U'\N{CJK UNIFIED IDEOGRAPH-2B740}' \
> || U'\U0002B81D' != U'\N{CJK UNIFIED IDEOGRAPH-2B81D}' \
> + || U'\U0002B81E' != U'\N{CJK UNIFIED IDEOGRAPH-2B81E}' \
> || U'\U0002B820' != U'\N{CJK UNIFIED IDEOGRAPH-2B820}' \
> || U'\U0002CEA1' != U'\N{CJK UNIFIED IDEOGRAPH-2CEA1}' \
> || U'\U0002CEB0' != U'\N{CJK UNIFIED IDEOGRAPH-2CEB0}' \
> @@ -175,10 +179,16 @@ typedef __CHAR32_TYPE__ char32_t;
> || U'\U000187F7' != U'\N{TANGUT IDEOGRAPH-187F7}' \
> || U'\U00018D00' != U'\N{TANGUT IDEOGRAPH-18D00}' \
> || U'\U00018D08' != U'\N{TANGUT IDEOGRAPH-18D08}' \
> + || U'\U00018D20' != U'\N{TANGUT IDEOGRAPH-18D20}' \
> || U'\U00018B00' != U'\N{KHITAN SMALL SCRIPT CHARACTER-18B00}' \
> || U'\U00018CD5' != U'\N{KHITAN SMALL SCRIPT CHARACTER-18CD5}' \
> + || U'\U00018CDA' != U'\N{KHITAN SMALL SCRIPT CHARACTER-18CDA}' \
> + || U'\U00018E00' != U'\N{JURCHEN CHARACTER-18E00}' \
> + || U'\U00019191' != U'\N{JURCHEN CHARACTER-19191}' \
> || U'\U0001B170' != U'\N{NUSHU CHARACTER-1B170}' \
> || U'\U0001B2FB' != U'\N{NUSHU CHARACTER-1B2FB}' \
> + || U'\U0003D000' != U'\N{SMALL SEAL CHARACTER-3D000}' \
> + || U'\U0003FC3F' != U'\N{SMALL SEAL CHARACTER-3FC3F}' \
> || U'\uF900' != U'\N{CJK COMPATIBILITY IDEOGRAPH-F900}' \
> || U'\uFA6D' != U'\N{CJK COMPATIBILITY IDEOGRAPH-FA6D}' \
> || U'\uFA70' != U'\N{CJK COMPATIBILITY IDEOGRAPH-FA70}' \
> --- a/libcpp/makeucnid.cc 2026-08-06 11:38:29.990867495 +0200
> +++ b/libcpp/makeucnid.cc 2026-09-17 10:12:45.327332655 +0200
> @@ -538,7 +538,7 @@ write_copyright (void)
> <http://www.gnu.org/licenses/>.\n\
> \n\
> \n\
> - Copyright (C) 1991-2025 Unicode, Inc. All rights reserved.\n\
> + Copyright (C) 1991-2026 Unicode, Inc. All rights reserved.\n\
> Distributed under the Terms of Use in\n\
> http://www.unicode.org/copyright.html.\n\
> \n\
> --- a/libcpp/makeuname2c.cc 2026-08-06 11:48:09.062090460 +0200
> +++ b/libcpp/makeuname2c.cc 2026-09-17 10:22:31.708306443 +0200
> @@ -69,7 +69,7 @@ struct entry { const char *name; unsigne
> static struct entry *entries;
> static unsigned long num_allocated, num_entries;
>
> -/* Unicode 16.0 Table 4-8. */
> +/* Unicode 18.0 Table 4-8. */
> struct generated {
> const char *prefix;
> /* max_high is a workaround for UnicodeData.txt inconsistencies
> @@ -84,7 +84,7 @@ static struct generated generated_ranges
> { "CJK UNIFIED IDEOGRAPH-", 0x4e00, 0x9fff, 0, 1, 0 },
> { "CJK UNIFIED IDEOGRAPH-", 0x20000, 0x2a6df, 0, 1, 0 },
> { "CJK UNIFIED IDEOGRAPH-", 0x2a700, 0x2b73f, 0, 1, 0 },
> - { "CJK UNIFIED IDEOGRAPH-", 0x2b740, 0x2b81d, 0, 1, 0 },
> + { "CJK UNIFIED IDEOGRAPH-", 0x2b740, 0x2b81e, 0, 1, 0 },
> { "CJK UNIFIED IDEOGRAPH-", 0x2b820, 0x2cead, 0, 1, 0 },
> { "CJK UNIFIED IDEOGRAPH-", 0x2ceb0, 0x2ebe0, 0, 1, 0 },
> { "CJK UNIFIED IDEOGRAPH-", 0x2ebf0, 0x2ee5d, 0, 1, 0 },
> @@ -93,12 +93,14 @@ static struct generated generated_ranges
> { "CJK UNIFIED IDEOGRAPH-", 0x323b0, 0x33479, 0, 1, 0 },
> { "EGYPTIAN HIEROGLYPH-", 0x13460, 0x143fa, 0, 2, 0 },
> { "TANGUT IDEOGRAPH-", 0x17000, 0x187ff, 0, 3, 0 },
> - { "TANGUT IDEOGRAPH-", 0x18d00, 0x18d1e, 0, 3, 0 },
> - { "KHITAN SMALL SCRIPT CHARACTER-", 0x18b00, 0x18cd5, 0, 4, 0 },
> - { "NUSHU CHARACTER-", 0x1b170, 0x1b2fb, 0, 5, 0 },
> - { "CJK COMPATIBILITY IDEOGRAPH-", 0xf900, 0xfa6d, 0, 6, 0 },
> - { "CJK COMPATIBILITY IDEOGRAPH-", 0xfa70, 0xfad9, 0, 6, 0 },
> - { "CJK COMPATIBILITY IDEOGRAPH-", 0x2f800, 0x2fa1d, 0, 6, 0 }
> + { "TANGUT IDEOGRAPH-", 0x18d00, 0x18d20, 0, 3, 0 },
> + { "KHITAN SMALL SCRIPT CHARACTER-", 0x18b00, 0x18cda, 0, 4, 0 },
> + { "JURCHEN CHARACTER-", 0x18e00, 0x19191, 0, 5, 0 },
> + { "NUSHU CHARACTER-", 0x1b170, 0x1b2fb, 0, 6, 0 },
> + { "SMALL SEAL CHARACTER-", 0x3d000, 0x3fc3f, 0, 7, 0 },
> + { "CJK COMPATIBILITY IDEOGRAPH-", 0xf900, 0xfa6d, 0, 8, 0 },
> + { "CJK COMPATIBILITY IDEOGRAPH-", 0xfa70, 0xfad9, 0, 8, 0 },
> + { "CJK COMPATIBILITY IDEOGRAPH-", 0x2f800, 0x2fa1d, 0, 8, 0 }
> };
>
> struct node {
> @@ -665,7 +667,7 @@ write_copyright (void)
> <http://www.gnu.org/licenses/>.\n\
> \n\
> \n\
> - Copyright (C) 1991-2025 Unicode, Inc. All rights reserved.\n\
> + Copyright (C) 1991-2026 Unicode, Inc. All rights reserved.\n\
> Distributed under the Terms of Use in\n\
> http://www.unicode.org/copyright.html.\n\
> \n\
> --- a/libstdc++-v3/include/bits/unicode.h 2026-05-30 17:45:09.527107977 +0200
> +++ b/libstdc++-v3/include/bits/unicode.h 2026-09-17 10:31:42.704765669 +0200
> @@ -753,9 +753,9 @@ namespace __unicode
> template<typename _View>
> using _Utf32_view = _Utf_view<char32_t, _View>;
>
> -inline namespace __v16_0_0
> +inline namespace __v18_0_0
> {
> -#define _GLIBCXX_GET_UNICODE_DATA 160000
> +#define _GLIBCXX_GET_UNICODE_DATA 180000
> #include "unicode-data.h"
> #ifdef _GLIBCXX_GET_UNICODE_DATA
> # error "Invalid unicode data"
> @@ -1118,7 +1118,7 @@ inline namespace __v16_0_0
> _Iterator _M_begin;
> };
>
> -} // namespace __v16_0_0
> +} // namespace __v18_0_0
>
> // Return the field width of a string.
> template<typename _CharT>
> --- a/libstdc++-v3/testsuite/ext/unicode/properties.cc 2026-03-27 10:17:23.566176826 +0100
> +++ b/libstdc++-v3/testsuite/ext/unicode/properties.cc 2026-09-17 11:11:47.320944575 +0200
> @@ -53,7 +53,7 @@ static_assert( uc::__field_width(U'\U000
> static_assert( uc::__field_width(U'\U0001FA69') == 1 );
> static_assert( uc::__field_width(U'\U0001FA70') == 2 );
> static_assert( uc::__field_width(U'\U0001FAF8') == 2 );
> -static_assert( uc::__field_width(U'\U0001FAF9') == 1 );
> +static_assert( uc::__field_width(U'\U0001FAF9') == 2 );
>
> using enum uc::_Gcb_property;
> static_assert( uc::__grapheme_cluster_break_property(U'\0') == _Gcb_Control );
>
> Jakub
More information about the Libstdc++
mailing list