This is the mail archive of the
gcc@gcc.gnu.org
mailing list for the GCC project.
Re: thoughts on martin's proposed patch for GCC and UTF-8
> Wouldn't it work for GCC to treat all byte values above 127 as part of
> an identifier, and not worry about how they group into multibyte
> characters?
That would work fine. It might or might not be what the user expects.
The only real drawback is standards compliance. C++, Java, and C9X all
allow to express Unicode in identifiers using \u escapes, like
void h\u00D6llo();
The standards go on saying that the actual source input might be in a
different character set, and the implementation defines how that
relates to Unicode. A sensible implementation would use translation
mechanisms.
Of course, the easiest thing would be to assume that we always get
non-ASCII in identifiers as Unicode escapes. The editor (e.g. Emacs)
would then need to convert the internal encoding to Unicode escapes
when saving. That would make the feature in the language truly useful.
Regards,
Martin