UTF-8 support - char or wchar_t?
Martin Sebor
sebor@roguewave.com
Mon Jun 21 17:26:00 GMT 2004
Ole Laursen wrote:
> Hi,
>
> Paolo Carlini <pcarlini@suse.de> writes:
>
>
>>Carlo Wood wrote:
>>
>>
>>>Can you please explain why one needs wchar_t (wcout)
>>>at all when using UTF-8? I'd expect that UTF-8 fits
>>>in a stream of 8-bit octets and no wchar_t should
>>>be needed at all.
>>>
>>
>>UTF-8 fits in a stream of 8-bit octets of the *external* encoding!
>>How do you produce this *external* UTF-8 encoded stream?
>>
>>The standard way is using an *internal* wchar_t representation
>>(basically, on GNU systems is UCS4, see the glibc docs), then
>>exploiting the specialization codecvt<wchar_t, char, mbstate_t>.
>
>
> So how do you do that? I currently use an ostringstream, how would I
> go about getting a std::string in the encoding of the current locale?
You can't. stringbuf doesn't use codecvt, only filebuf does. It's
kind of a pain. One way to do it is to write the data to a file
(or a socket) using an ofstream (or wofstream) with your locale
imbued in it, then read it back in using a basic_ifstream<char>
with the "C" locale imbued in it. The more efficient but more
cumbersome way to do it (w/o file I/O) is to use the codecvt
facet directly. The most portable way to do this, though, is
to forget about codecvt and use iconv directly. Attached is
a simple wrapper function I once wrote for someone who didn't
like the clunky iconv interface. Read the pages below for more
info on iconv:
http://www.opengroup.org/onlinepubs/009695399/functions/iconv.html
Martin
-------------- next part --------------
An embedded and charset-unspecified text was scrubbed...
Name: iconv.cpp
URL: <http://gcc.gnu.org/pipermail/libstdc++/attachments/20040621/47b04dcd/attachment.ksh>
More information about the Libstdc++
mailing list