[PATCH 1/3] contrib, libcpp, libstdc++: Update to Unicode 18.0.0

Jonathan Wakely jwakely@redhat.com
Thu Sep 17 10:56:31 GMT 2026


On Thu, 17 Sept 2026 at 11:04, Jakub Jelinek <jakub@redhat.com> wrote:
>
> Hi!
>
> Unicode 18.0.0 has been released yesterday.
> The following patch follows the unicode/README and updates
> whatever is needed in gcc to that Unicode version.  In glibc there
> weren't any changes to the from_glibc files.
>
> Note, apparently I forgot last year to bump in libstdc++ the
> inline namespace and 160000 code to 170000, this bumps it right to
> 180000.  As for tests, I've added some new characters to
> named-universal-char-escape-1.c and made sure to test boundaries of
> generated ranges where it changed and test new generated ranges.
> In libstdc++ testsuite/ext/unicode/properties.cc test I had to tweak
> one static assert because EastAsianWidth.txt changed for the character
> being tested (and there is no character right above U+1FAFA).
>
> As usually, the patch is way too large for the mailing list limits
> even when xz -9e compressed.
> Below is included the ChangeLog and the non-generated changes, attached
> is xz -9e compressed first part of the generated changes (all
> generated files except the huge uname2c.h second table), in a follow-up
> I'll post one xz -9e compressed patch which removes the uname2c.h
> second table and another one which adds it again with the new content
> (basically the whole array is changed, so a diff which replaces it
> is alone also too large).
>
> So far lightly tested, ok for trunk if it passes full bootstrap/regtest?

The libstdc++ parts are OK, thanks.


>
> 2026-09-17  Jakub Jelinek  <jakub@redhat.com>
>
> contrib/
>         * unicode/README: Replace Unicode 17 version with 18.
>         * unicode/gen_libstdcxx_unicode_data.py: Use 180000 instead of
>         160000 in _GLIBCXX_GET_UNICODE_DATA test.
>         * unicode/DerivedCoreProperties.txt: Updated from Unicode 18.0.
>         * unicode/emoji-data.txt: Likewise.
>         * unicode/PropList.txt: Likewise.
>         * unicode/HangulSyllableType.txt: Likewise.
>         * unicode/DerivedNormalizationProps.txt: Likewise.
>         * unicode/NameAliases.txt: Likewise.
>         * unicode/UnicodeData.txt: Likewise.
>         * unicode/EastAsianWidth.txt: Likewise.
> gcc/testsuite/
>         * c-c++-common/cpp/named-universal-char-escape-1.c: Add tests
>         for some Unicode 18.0 characters, both normal and generated.
> libcpp/
>         * makeucnid.cc (write_copyright): Update Unicode Copyright years.
>         * makeuname2c.cc (generated_ranges): Adjust Unicode version from 17.0
>         to 18.0.  Add JURCHEN CHARACTER- and SMALL SEAL CHARACTER- generated
>         ranges, adjust indexes in following entries.
>         (write_copyright): Update Unicode Copyright years.
>         * generated_cpp_wcwidth.h: Regenerated.
>         * ucnid.h: Regenerated.
>         * uname2c.h: Regenerated.
> libstdc++-v3/
>         * include/bits/unicode.h (std::__unicode::__v16_0_0): Rename inline
>         namespace to ...
>         (std::__unicode::__v18_0_0): ... this.
>         (_GLIBCXX_GET_UNICODE_DATA): Change from 160000 to 180000.
>         * testsuite/ext/unicode/properties.cc: Expect 2 rather than 1 for
>         uc::__field_width(U'\U0001FAF9').
>         * include/bits/unicode-data.h: Regenerated.
>
> --- a/contrib/unicode/README    2026-08-06 11:38:29.988537277 +0200
> +++ b/contrib/unicode/README    2026-09-17 10:10:21.815296372 +0200
> @@ -47,14 +47,14 @@ produce ucnid.h.
>
>  The procedure to update GCC's Unicode support is the following:
>
> -1.  Update the six Unicode data files from the above URLs.
> +1.  Update the ten Unicode data files from the above URLs.
>
>  2.  Update the two glibc files in from_glibc/ from glibc's git.  Update
>      the commit number above in this README.
>
>  3.  Run ./gen_wcwidth.py X.Y > ../../libcpp/generated_cpp_wcwidth.h
>      (where X.Y is the version of the Unicode standard corresponding to the
> -    Unicode data files being used, most recently, 17.0.0).
> +    Unicode data files being used, most recently, 18.0.0).
>
>  4.  Update Unicode Copyright years in libcpp/makeucnid.cc and in
>      libcpp/makeuname2c.cc up to the year in which the Unicode
> @@ -69,7 +69,7 @@ The procedure to update GCC's Unicode su
>         PropList.txt > ../../libcpp/ucnid.h
>
>  7.  Read the corresponding Unicode's standard and update correspondingly
> -    generated_ranges table in libcpp/makeuname2c.cc (in Unicode 17 all
> +    generated_ranges table in libcpp/makeuname2c.cc (in Unicode 18 all
>      the needed information was in Table 4-8).
>
>  8.  Compile makeuname2c, e.g. with:
> --- a/contrib/unicode/gen_libstdcxx_unicode_data.py     2026-03-27 10:17:13.244345262 +0100
> +++ b/contrib/unicode/gen_libstdcxx_unicode_data.py     2026-09-17 10:29:44.627380770 +0200
> @@ -64,7 +64,7 @@ print("""
>  """)
>  print("#ifndef _GLIBCXX_GET_UNICODE_DATA")
>  print('# error "This is not a public header, do not include it directly"')
> -print("#elif _GLIBCXX_GET_UNICODE_DATA != 160000")
> +print("#elif _GLIBCXX_GET_UNICODE_DATA != 180000")
>  print('# error "Version mismatch for Unicode static data"')
>  print("#endif\n")
>
> --- a/gcc/testsuite/c-c++-common/cpp/named-universal-char-escape-1.c    2026-03-27 10:17:15.174313768 +0100
> +++ b/gcc/testsuite/c-c++-common/cpp/named-universal-char-escape-1.c    2026-09-17 10:59:47.165738592 +0200
> @@ -119,6 +119,9 @@ typedef __CHAR32_TYPE__ char32_t;
>      || U'\uFE18' != U'\N{PRESENTATION FORM FOR VERTICAL RIGHT WHITE LENTICULAR BRACKET}' \
>      || U'\u0CF3' != U'\N{KANNADA SIGN COMBINING ANUSVARA ABOVE RIGHT}' \
>      || U'\u0ECE' != U'\N{LAO YAMAKKAN}' \
> +    || U'\u05C9' != U'\N{HEBREW POINT DAGESH HAZAQ MUDGASH}' \
> +    || U'\u20C3' != U'\N{UAE DIRHAM SIGN}' \
> +    || U'\U00010EE1' != U'\N{ARABIC CROWN LETTER SHEEN}' \
>      || U'\U00010EFE' != U'\N{ARABIC SMALL LOW WORD QASR}' \
>      || U'\U00011241' != U'\N{KHOJKI VOWEL SIGN VOCALIC R}' \
>      || U'\U00011B06' != U'\N{DEVANAGARI SIGN WESTERN FIVE-LIKE BHALE}' \
> @@ -160,6 +163,7 @@ typedef __CHAR32_TYPE__ char32_t;
>      || U'\U0002B739' != U'\N{CJK UNIFIED IDEOGRAPH-2B739}' \
>      || U'\U0002B740' != U'\N{CJK UNIFIED IDEOGRAPH-2B740}' \
>      || U'\U0002B81D' != U'\N{CJK UNIFIED IDEOGRAPH-2B81D}' \
> +    || U'\U0002B81E' != U'\N{CJK UNIFIED IDEOGRAPH-2B81E}' \
>      || U'\U0002B820' != U'\N{CJK UNIFIED IDEOGRAPH-2B820}' \
>      || U'\U0002CEA1' != U'\N{CJK UNIFIED IDEOGRAPH-2CEA1}' \
>      || U'\U0002CEB0' != U'\N{CJK UNIFIED IDEOGRAPH-2CEB0}' \
> @@ -175,10 +179,16 @@ typedef __CHAR32_TYPE__ char32_t;
>      || U'\U000187F7' != U'\N{TANGUT IDEOGRAPH-187F7}' \
>      || U'\U00018D00' != U'\N{TANGUT IDEOGRAPH-18D00}' \
>      || U'\U00018D08' != U'\N{TANGUT IDEOGRAPH-18D08}' \
> +    || U'\U00018D20' != U'\N{TANGUT IDEOGRAPH-18D20}' \
>      || U'\U00018B00' != U'\N{KHITAN SMALL SCRIPT CHARACTER-18B00}' \
>      || U'\U00018CD5' != U'\N{KHITAN SMALL SCRIPT CHARACTER-18CD5}' \
> +    || U'\U00018CDA' != U'\N{KHITAN SMALL SCRIPT CHARACTER-18CDA}' \
> +    || U'\U00018E00' != U'\N{JURCHEN CHARACTER-18E00}' \
> +    || U'\U00019191' != U'\N{JURCHEN CHARACTER-19191}' \
>      || U'\U0001B170' != U'\N{NUSHU CHARACTER-1B170}' \
>      || U'\U0001B2FB' != U'\N{NUSHU CHARACTER-1B2FB}' \
> +    || U'\U0003D000' != U'\N{SMALL SEAL CHARACTER-3D000}' \
> +    || U'\U0003FC3F' != U'\N{SMALL SEAL CHARACTER-3FC3F}' \
>      || U'\uF900' != U'\N{CJK COMPATIBILITY IDEOGRAPH-F900}' \
>      || U'\uFA6D' != U'\N{CJK COMPATIBILITY IDEOGRAPH-FA6D}' \
>      || U'\uFA70' != U'\N{CJK COMPATIBILITY IDEOGRAPH-FA70}' \
> --- a/libcpp/makeucnid.cc       2026-08-06 11:38:29.990867495 +0200
> +++ b/libcpp/makeucnid.cc       2026-09-17 10:12:45.327332655 +0200
> @@ -538,7 +538,7 @@ write_copyright (void)
>     <http://www.gnu.org/licenses/>.\n\
>  \n\
>  \n\
> -   Copyright (C) 1991-2025 Unicode, Inc.  All rights reserved.\n\
> +   Copyright (C) 1991-2026 Unicode, Inc.  All rights reserved.\n\
>     Distributed under the Terms of Use in\n\
>     http://www.unicode.org/copyright.html.\n\
>  \n\
> --- a/libcpp/makeuname2c.cc     2026-08-06 11:48:09.062090460 +0200
> +++ b/libcpp/makeuname2c.cc     2026-09-17 10:22:31.708306443 +0200
> @@ -69,7 +69,7 @@ struct entry { const char *name; unsigne
>  static struct entry *entries;
>  static unsigned long num_allocated, num_entries;
>
> -/* Unicode 16.0 Table 4-8.  */
> +/* Unicode 18.0 Table 4-8.  */
>  struct generated {
>    const char *prefix;
>    /* max_high is a workaround for UnicodeData.txt inconsistencies
> @@ -84,7 +84,7 @@ static struct generated generated_ranges
>    { "CJK UNIFIED IDEOGRAPH-", 0x4e00, 0x9fff, 0, 1, 0 },
>    { "CJK UNIFIED IDEOGRAPH-", 0x20000, 0x2a6df, 0, 1, 0 },
>    { "CJK UNIFIED IDEOGRAPH-", 0x2a700, 0x2b73f, 0, 1, 0 },
> -  { "CJK UNIFIED IDEOGRAPH-", 0x2b740, 0x2b81d, 0, 1, 0 },
> +  { "CJK UNIFIED IDEOGRAPH-", 0x2b740, 0x2b81e, 0, 1, 0 },
>    { "CJK UNIFIED IDEOGRAPH-", 0x2b820, 0x2cead, 0, 1, 0 },
>    { "CJK UNIFIED IDEOGRAPH-", 0x2ceb0, 0x2ebe0, 0, 1, 0 },
>    { "CJK UNIFIED IDEOGRAPH-", 0x2ebf0, 0x2ee5d, 0, 1, 0 },
> @@ -93,12 +93,14 @@ static struct generated generated_ranges
>    { "CJK UNIFIED IDEOGRAPH-", 0x323b0, 0x33479, 0, 1, 0 },
>    { "EGYPTIAN HIEROGLYPH-", 0x13460, 0x143fa, 0, 2, 0 },
>    { "TANGUT IDEOGRAPH-", 0x17000, 0x187ff, 0, 3, 0 },
> -  { "TANGUT IDEOGRAPH-", 0x18d00, 0x18d1e, 0, 3, 0 },
> -  { "KHITAN SMALL SCRIPT CHARACTER-", 0x18b00, 0x18cd5, 0, 4, 0 },
> -  { "NUSHU CHARACTER-", 0x1b170, 0x1b2fb, 0, 5, 0 },
> -  { "CJK COMPATIBILITY IDEOGRAPH-", 0xf900, 0xfa6d, 0, 6, 0 },
> -  { "CJK COMPATIBILITY IDEOGRAPH-", 0xfa70, 0xfad9, 0, 6, 0 },
> -  { "CJK COMPATIBILITY IDEOGRAPH-", 0x2f800, 0x2fa1d, 0, 6, 0 }
> +  { "TANGUT IDEOGRAPH-", 0x18d00, 0x18d20, 0, 3, 0 },
> +  { "KHITAN SMALL SCRIPT CHARACTER-", 0x18b00, 0x18cda, 0, 4, 0 },
> +  { "JURCHEN CHARACTER-", 0x18e00, 0x19191, 0, 5, 0 },
> +  { "NUSHU CHARACTER-", 0x1b170, 0x1b2fb, 0, 6, 0 },
> +  { "SMALL SEAL CHARACTER-", 0x3d000, 0x3fc3f, 0, 7, 0 },
> +  { "CJK COMPATIBILITY IDEOGRAPH-", 0xf900, 0xfa6d, 0, 8, 0 },
> +  { "CJK COMPATIBILITY IDEOGRAPH-", 0xfa70, 0xfad9, 0, 8, 0 },
> +  { "CJK COMPATIBILITY IDEOGRAPH-", 0x2f800, 0x2fa1d, 0, 8, 0 }
>  };
>
>  struct node {
> @@ -665,7 +667,7 @@ write_copyright (void)
>     <http://www.gnu.org/licenses/>.\n\
>  \n\
>  \n\
> -   Copyright (C) 1991-2025 Unicode, Inc.  All rights reserved.\n\
> +   Copyright (C) 1991-2026 Unicode, Inc.  All rights reserved.\n\
>     Distributed under the Terms of Use in\n\
>     http://www.unicode.org/copyright.html.\n\
>  \n\
> --- a/libstdc++-v3/include/bits/unicode.h       2026-05-30 17:45:09.527107977 +0200
> +++ b/libstdc++-v3/include/bits/unicode.h       2026-09-17 10:31:42.704765669 +0200
> @@ -753,9 +753,9 @@ namespace __unicode
>    template<typename _View>
>      using _Utf32_view = _Utf_view<char32_t, _View>;
>
> -inline namespace __v16_0_0
> +inline namespace __v18_0_0
>  {
> -#define _GLIBCXX_GET_UNICODE_DATA 160000
> +#define _GLIBCXX_GET_UNICODE_DATA 180000
>  #include "unicode-data.h"
>  #ifdef _GLIBCXX_GET_UNICODE_DATA
>  # error "Invalid unicode data"
> @@ -1118,7 +1118,7 @@ inline namespace __v16_0_0
>        _Iterator _M_begin;
>      };
>
> -} // namespace __v16_0_0
> +} // namespace __v18_0_0
>
>    // Return the field width of a string.
>    template<typename _CharT>
> --- a/libstdc++-v3/testsuite/ext/unicode/properties.cc  2026-03-27 10:17:23.566176826 +0100
> +++ b/libstdc++-v3/testsuite/ext/unicode/properties.cc  2026-09-17 11:11:47.320944575 +0200
> @@ -53,7 +53,7 @@ static_assert( uc::__field_width(U'\U000
>  static_assert( uc::__field_width(U'\U0001FA69') == 1 );
>  static_assert( uc::__field_width(U'\U0001FA70') == 2 );
>  static_assert( uc::__field_width(U'\U0001FAF8') == 2 );
> -static_assert( uc::__field_width(U'\U0001FAF9') == 1 );
> +static_assert( uc::__field_width(U'\U0001FAF9') == 2 );
>
>  using enum uc::_Gcb_property;
>  static_assert( uc::__grapheme_cluster_break_property(U'\0') == _Gcb_Control );
>
>         Jakub



More information about the Libstdc++ mailing list