arch-specific template code

Paolo Carlini paolo.carlini@oracle.com
Mon Aug 27 22:27:00 GMT 2012


Hi,

On 08/27/2012 04:27 PM, Ulrich Drepper wrote:
> The next question is how to to implement arch-specific optimizations
> to template code.
>
> Obviously, these headers have to be multilib aware and therefore
> handle, e.g., x86, x86-64, and perhaps x32 concurrently.  But having
> code for unrelated architectures in the same file would probably be
> rather distracting.
>
> There will likely be two different situations:
> - entire functions are replaced
> - specializations are provided
>
> The former probably will require constructs of the form
>
>    #ifndef _SOME_FUNCTION_defined
>    ...
>    #endif
>
> Should these files go into libstdc++-v3/config/cpu ?
First blush, it seems the right place, yes.
> Aside from <random> there will be a few more headers which can benefit
> from such optimizations.
>
> Any opinions?
My personal opinion is that a concrete example, small, but meaningful 
and rather self contained, would help. To be honest, at this stage, 
isn't clear to me which kind of arch-specific optimizations you are 
thinking about. For example, I mentioned in conversation tons of times 
that it would probably make sense delivering a <numeric> explicitly 
exploiting vectorization (thus much beyond what the autovectorizer can 
do). Do you have in mind something similar? Another very vague idea of 
mind, would be playing with some "easy" containers, like <forward_list> 
and TM, like internally wrapping some operations in transactions (this 
may become more concrete when many of us have the hardware of course).

In short, personally, I would suggest discussing general design ideas 
starting anyway from something rather concrete.

Paolo.



More information about the Libstdc++ mailing list