This is the mail archive of the fortran@gcc.gnu.org mailing list for the GNU Fortran project.
| Index Nav: | [Date Index] [Subject Index] [Author Index] [Thread Index] | |
|---|---|---|
| Message Nav: | [Date Prev] [Date Next] | [Thread Prev] [Thread Next] |
| Other format: | [Raw text] | |
I was recently looking over the fortran Sweep3D benchmark and noticed that
in the inner most loops it was making heavy use of the SIGN intrinsic.
Looking at the gimple we generate for the integer variants of the
gfortran's SIGN intrinsic, we can not only do slightly better but at the
same time fix a latent bug.
The result of SIGN(X,Y) is the absolute value of X but with the sign of Y.
For floating point types this conveniently maps to __builtin_copysign, but
for integral types we currently lower this to the tree:
return (a >= 0) ^ (b >= 0) ? -a : a
which on x86_64 expands to something like:
movl %edi, %eax
notl %esi
movl %edi, %edx
notl %eax
shrl $31, %esi
negl %edx
shrl $31, %eax
xorl %esi, %eax
cmovne %edx, %edi
movl %edi, %eax
ret
It turns out that its possible to do this more conveniently without
branching or conditional moves using the alternate sequence:
int tmp = (a ^ b) >> 31;
return (a + t) ^ t;
The keen eyed will recognize this as very similar to the branchless "abs"
that we generate in the middle-end, but tweaked such that we perform the
negation if the sign bits differ. Interested readers can also find this
implementation suggested in section 2-9 "Transfer of Sign" in Henry
Warren's book "Hacker's Delight".
With this alternate implementation, we now generate the shorter and faster:
xorl %edi, %esi
sarl $31, %esi
leal (%rsi,%rdi), %eax
xorl %esi, %eax
ret
I tried quantifying how much better this was by comparing the times to call
the above two implementations several million times vs. a control of
calling a function that just returned zero. Unfortunately, this approach
is severely flawed with the short sequence overlapping considerably with
the call/loop overhead; the resulting times were 20.949s for a null
implementation, 20.954 for the new implementation and 24.586 for the old
implementation resulting in a apparent speed-up of several hundred fold!?
In theory, many of the above transformations should be implemented in
fold-const.c and the various if-conversion passes (and I'll submit those
patches as follow-ups), but whilst investigating this I noticed that
there's currently a bug in the current SIGN intrinsic expansion, in that
it doesn't protect against multiple evaluation of it's arguments.
So the FORTRAN code:
function foo()
integer :: foo
foo = sign(ia(),ib())
end function
end
on mainline expands to the tree
foo ()
{
int4 __result_foo;
__result_foo = ia () >= 0 ^ ib () >= 0 ? -ia () : ia ();
return __result_foo;
}
where we clearly call the function ia(), the X argument, multiple times.
When this function has side-effects, we can end up generating incorrect
results. In the middle-end, we'd normally use a SAVE_EXPR to handle this,
but the gfortran scalarizer conveniently provides a gfc_evaluate_now
function for this purpose. If I could write FORTRAN, I might even have
attempted a new testcase for the testsuite.
The following patch has been tested on x86_64-unknown-linux-gnu with a
full "make bootstrap", all default languages, and regression tested with a
top-level "make -k check" with no new failures.
Ok for mainline?
2007-01-19 Roger Sayle <roger@eyesopen.com>
* trans-intrinsic.c (gfc_conv_intrinsic_sign): New branchless
implementation for the SIGN intrinsic with integral operands.
(gfc_conv_intrinsic_minmax): Fix whitespace.
Roger
--
Attachment:
patchs.txt
Description: Text document
| Index Nav: | [Date Index] [Subject Index] [Author Index] [Thread Index] | |
|---|---|---|
| Message Nav: | [Date Prev] [Date Next] | [Thread Prev] [Thread Next] |