This is the mail archive of the
gcc@gcc.gnu.org
mailing list for the GCC project.
Re: thoughts on martin's proposed patch for GCC and UTF-8
Date: Thu, 10 Dec 1998 08:12:20 +0100
From: Martin von Loewis <martin@mira.isdn.cs.tu-berlin.de>
Also, if the mangling of gxxint.texi is used, Fo\u1234 becomes
U7Fo_1234, where the U indicates that the underscore is an escape....
Sorry, I'm still lost. If the identifier is the UTF-8 character MICRO
SIGN (code 00B5), do you generate the same UTF-8 character on output,
or do you mangle it as if the user had typed `\u00b5'? If the latter,
then I don't understand why gas needs to be 8-bit clean; if the
former, then I don't understand your example with \u1234 as it seems
to me that it won't unify with the UTF-8 sequence that is equivalent
to \u1234.
> I've run into shells that use the top bit for their own purposes.
What system?
Older BSD systems. The original Bourne shell used the top bit for its
own purposes. A few years back, all major Unix suppliers went through
their shells and made them 8-bit clean, but a few bugs lurked for a
while and I wouldn't be surprised if some were still out there.
Please tell me how I can perform the same test with ASCII-only
shell commands, and I happily convert.
Something like this should do it:
echo ab | tr 'ab' '\123\456'
Or you could write a little C program, compile it, and run it.
> C++ does not distinguish between non-ASCII digits and letters.
>
> Really? Suppose I write the preprocessor line
>
> #if X == 1
>
> where X is some Japanese identifier, but I make the understandable
> mistake of using a FULLWIDTH DIGIT ONE (code FF11) instead of an ASCII 1.
\uFF11 is not a letter in C++
OK, so then there's no problem: C++ _does_ distinguish between
non-ASCII digits and letters.