optimizations

Andrew Pinski pinskia@physics.uc.edu
Tue Feb 18 18:16:00 GMT 2003


Here is the generation for 3.4 (20030214) on ppc (darwin/mac os x):
_main:
         mflr r2
         li r9,0
         stw r2,8(r1)
         li r2,0
         stwu r1,-64(r1)
         stw r2,56(r1)		<== store 0 in &k before the loop. WHY?
         b L2
L10:
         addi r9,r9,1
L2:
         cmpwi cr0,r9,16
         bne+ cr0,L10
         addi r4,r1,56	<== the second parm to write
         li r3,1		<== the first parm to write
         li r5,1		<== the third parm to write
         stw r9,56(r1)	<== stores the 2nd parm to write aka &k
         bl _write		<== `call' write
         addi r1,r1,64	<== restore the stack pointer
         lwz r4,8(r1)	
         li r3,0		<== the return value from main
         mtlr r4
         blr

Here is the generation for 3.4 (20030215) on x86 (linux):

main:
         pushl   %ebp
         movl    %esp, %ebp
         subl    $24, %esp
         movl    $0, -4(%ebp)	<=== store 0 into &k
         andl    $-16, %esp
         jmp     .L2
         .p2align 4,,7
.L9:
         incl    %eax			<=== increment k
         movl    %eax, -4(%ebp)	<=== store k into &k WHY?
.L2:
         movl    -4(%ebp), %eax	<=== load k from &k WHY?
         cmpl    $16, %eax		<=== compare k to 16
         jne     .L9			<=== jump not equal to L9
         movl    $1, 8(%esp)		<=== 1st parm to write (1)
         leal    -4(%ebp), %edx	<=== (&k) into edx
         movl    %edx, 4(%esp)	<=== 2nd parm to write (&k/edx)
         movl    $1, (%esp)		<=== 3rd parm to write (1)
         call    write			<=== call write
         leave
         ret

 From the looks of it, gcc does a better job on ppc compared to i686 for 
some reason in terms of optimizations.


Thanks,
Andrew Pinski



On Tuesday, Feb 18, 2003, at 09:55 US/Pacific, Håkan Hjort wrote:

> Wed Jan 15 2003, Reza Roboubi wrote:
>> tm_gccmail@mail.kloo.net wrote:
>> [snap]
>>> I mentioned this on the gcc-bugs mailing list, and Mark Mitchell
>>> contributed a fairly simple load hoisting improvement to the loop
>>> optmiizer which restored performance on Whetstone.
>>>
>>> If you look at the gcc-bugs archives for 1998, you may be able to 
>>> find
>>> this message thread.
>> [snap]
>>
>> Thanks for this input. It would be interesting to see how the issue 
>> was fixed.
> Sorry for getting into this so late.
> Nobody actually posted the code generated by 3.3/3.4...
>
> inline int mm(int *i) {
>         if((*i)==0x10) return 0;
>         (*i)++; return 1;
> }
>
> int main() {
>         int k=0;
>         while (mm(&k)) {}
>         write(1,&k,1);
>         return 0;
> }
>
> For Sun's Forte compiler one gets the following:
>
> main:
>          save    %sp,-104,%sp
>          or      %g0,16,%g1
>          st      %g1,[%fp-4]
>          add     %fp,-4,%o1
>          or      %g0,1,%o0
>          call    write   ! params =  %o0 %o1 %o2 ! Result
>          or      %g0,1,%o2
>          ret     ! Result =  %i0
>          restore %g0,0,%o0
>
> I.e. it just stores '16' in k before the call to write, no trace left
> of mm() or any loop, as should be.
>
> Perhaps GCC now does the same after hoisting both the load and the 
> store?
>
> -- 
> /Håkan
>



More information about the Gcc mailing list