arch-specific template code
Paolo Carlini
paolo.carlini@oracle.com
Mon Aug 27 22:27:00 GMT 2012
Hi,
On 08/27/2012 04:27 PM, Ulrich Drepper wrote:
> The next question is how to to implement arch-specific optimizations
> to template code.
>
> Obviously, these headers have to be multilib aware and therefore
> handle, e.g., x86, x86-64, and perhaps x32 concurrently. But having
> code for unrelated architectures in the same file would probably be
> rather distracting.
>
> There will likely be two different situations:
> - entire functions are replaced
> - specializations are provided
>
> The former probably will require constructs of the form
>
> #ifndef _SOME_FUNCTION_defined
> ...
> #endif
>
> Should these files go into libstdc++-v3/config/cpu ?
First blush, it seems the right place, yes.
> Aside from <random> there will be a few more headers which can benefit
> from such optimizations.
>
> Any opinions?
My personal opinion is that a concrete example, small, but meaningful
and rather self contained, would help. To be honest, at this stage,
isn't clear to me which kind of arch-specific optimizations you are
thinking about. For example, I mentioned in conversation tons of times
that it would probably make sense delivering a <numeric> explicitly
exploiting vectorization (thus much beyond what the autovectorizer can
do). Do you have in mind something similar? Another very vague idea of
mind, would be playing with some "easy" containers, like <forward_list>
and TM, like internally wrapping some operations in transactions (this
may become more concrete when many of us have the hardware of course).
In short, personally, I would suggest discussing general design ideas
starting anyway from something rather concrete.
Paolo.
More information about the Libstdc++
mailing list