Major deficiency in optimising memset calls
Daniel Egger
degger@fhm.edu
Thu Oct 14 16:07:00 GMT 2004
Hija,
I'm currently trying to track down why gcc is generating
this poor code when optimising memset with constant params.
A helping hand would be greatly appreciated.
Consider the following code:
char a[8]
int main (void)
{
__builtin_memset (a, 0, 8);
}
This will result in two simple stores of the value 0 into
two adjacent 32bit values in memory on at least darwin-powerpc
and linux-i686. Great!
Next try:
char a[8]
int main (void)
{
__builtin_memset (a, 1, 8);
}
will result in:
pushl %ebp
movl $16843009, %eax
movl %esp, %ebp
subl $8, %esp
movl $16843009, %edx
andl $-16, %esp
movl %eax, a+4
subl $16, %esp
movl %edx, a
leave
ret
on linux-i686 compiled with gcc-3.4 -c -march=i686 -mtune=i686 -O2.
Let's call that in interesting result.
gcc-4.0 (todays CVS) on darwin completely looses track here:
_main:
lis r3,ha16(L_a$non_lazy_ptr)
li r4,1
li r5,8
lwz r3,lo16(L_a$non_lazy_ptr)(r3)
b L_memset$stub
.comm _a,8
.data
.section __TEXT,__symbol_stub1,symbol_stubs,pure_instructions,16
.align 4
L_memset$stub:
.indirect_symbol _memset
lis r11,ha16(L_memset$lazy_ptr)
lwzu r12,lo16(L_memset$lazy_ptr)(r11)
mtctr r12
bctr
.data
.lazy_symbol_pointer
L_memset$lazy_ptr:
.indirect_symbol _memset
.long dyld_stub_binding_helper
.data
.non_lazy_symbol_pointer
L_a$non_lazy_ptr:
.indirect_symbol _a
.long 0
.subsections_via_symbols
And it keeps getting worse:
char a[12]
int main (void)
{
__builtin_memset (a, 0, 12);
}
aforementioned gcc call on i686 will result in:
main:
pushl %ebp
xorl %eax, %eax
movl %esp, %ebp
pushl %edi
cld
subl $4, %esp
movl $a, %edi
movl $3, %ecx
andl $-16, %esp
subl $16, %esp
rep
stosl
movl -4(%ebp), %edi
leave
ret
The rep;stosl sequence seems somewhat okay to me for a simply
inline sequence to clear a larger or only at runtime known
amount of memory, however for under 16 bytes this seems like a
clear loss in performance and code size to me.
On darwin-powerpc we get the same sequence as above except for
the to-be-written value and the amount which is a clear lose
no matter what in this case.
There are several facts known in this case which should help
to generate much better code:
- The allocated memory is local and not volatile
- The destination of the store is unused
- The alignment and size of the storage
- The destination memory is a (possibly large) block instead
of scattered locations
- The parameters to the memset call are all constant
In addition to the poor code on i686 the results on darwin-powerpc
are really abysmal because we there seems to be no way to force
the inlining of memset which always means a *very* costly rountrip
into the linkage even for obvious uses.
I'm very certain that this pessimizes quite a lot of common
applications (including gcc) where a memset is chosen in favor
of a loop to clear a block of memory.
I'm currently digging into builtins.c but a helping hand would
be more than appreciated.
Servus,
Daniel
-------------- next part --------------
A non-text attachment was scrubbed...
Name: PGP.sig
Type: application/pgp-signature
Size: 478 bytes
Desc: This is a digitally signed message part
URL: <https://gcc.gnu.org/pipermail/gcc/attachments/20041014/0332dc14/attachment.sig>
More information about the Gcc
mailing list