This is the mail archive of the
gcc@gcc.gnu.org
mailing list for the GCC project.
Re: thoughts on martin's proposed patch for GCC and UTF-8
By the way, even if we don't care about linking from different
locales, GCC must still translate symbols to a canonical form. For
example, suppose `@' denotes the character MICRO SIGN (Unicode
character 00b5). Then `@' (1 character) and `\u00b5' (6 characters)
are different spellings of the same symbol, and GCC must unify the two
spellings. This is true no matter how the symbol is represented in
assembly language output.
That's right: GCC will have to convert \u00b5 into whatever is
the proper thing to output for @, and likewise for unicode characters
that have a multibyte representation in the encoding system it is using.
This is why GCC has to depend on the encoding, when \u is used.