RFC: fp printing speedup patch

Paolo Carlini pcarlini@unitus.it
Thu Nov 13 09:18:00 GMT 2003


Jerry Quinn wrote:

>I agree that a loop must be used instead of an array call.
>
Ok, this is a trivial change.

>  My point
>was that we don't want do be doing widen() at all during _M_insert_int
>or float.  At the moment for int, the widen call happens when
>numpunct_cache is constructed.
>
Ok, now I see, of course you are right sorry.

> > Perhaps because of what I wrote above? If you have to properly convert
> > all the chars anyway, in the most common case (plain chars no user
> > provided do_widen) many virtual function calls are still faster than
> > memcopy.
>
>Why?  A memcopy is a single call, and many virtual function calls are
>going to cost much more by comparison.  The work happening in these
>functions is small, which is why it's an issue.
>
My point was that a few virtul function calls doing nop should (and in 
fact if
you benchmark _M_insert_float this is the case) not cost more that one
function call doing memcopy inside, however.

>  However, the numpunct
>cache should amortize the cost to nothing for a real program.
>
Ok, now everything is clear to me.

>The bug is that when the new locale is created, the array in
>__numpunct_cache doesn't get updated by redoing the widen calls on
>this array with the new ctype facet.  As a result, we get integers
>rendered using the standard ctype instead of the modified one.
>
Ah! Will add some comments to my PR pointing out that the real reason
of the trouble is completely different.

>I went back and figured out that I had run 3.4 on the unaltered
>library, i.e. I forgot to install the modified library.
>
>So, here are the updated numbers.  I'm actually very surprised!  I
>didn't think we could do better than 2.95 since 2.95 didn't do a
>number of the things required by the standard:
>
>2.95
>
>real    0m4.383s
>user    0m4.380s
>sys     0m0.000s
>
>3.3.2
>
>real    0m14.521s
>user    0m14.310s
>sys     0m0.010s
>
>3.4
>
>real    0m7.026s
>user    0m7.010s
>sys     0m0.000s
>
>3.4+fppatch
>
>real    0m2.884s
>user    0m2.880s
>sys     0m0.000s
>
Great! Now everything is consistent, or even *better* than what could be
expected on the basis of the initial analysis. At this point I would be
really curious to know what happens for your testcase involving big numbers
too: probably that whole picture (low level output included) is quite
good for those too!

Paolo.



More information about the Libstdc++ mailing list