rtlopt loop unroller question
Zdenek Dvorak
rakdver@atrey.karlin.mff.cuni.cz
Wed Oct 22 01:08:00 GMT 2003
Hello,
> The following (sent on behalf of Yossi Markovich) simple loop:
>
> int * foo ()
> {
> int A[N];
> int B[N];
> int i;
> for (i=0; i<N; i++)
> A[i] = B[i];
> return A;
> }
>
> results in much better code when compiled using "gcc3.4 -O3
> -fold-unroll-loops", than when compiled using the rltopt branch with "-O3
> -funroll-loops" (on powerpc-apple-darwin6.4). We are aware of the fact that
> the new loop optimizer in mainline is known to have caused regressions; we
> were wondering whether something can be done to get the better addressing
> calculation using the rltopt branch (possibly using a different set of
> flags?)?
rtlopt branch contains a completely new loop optimizer, which is quite
experimental (and I never got to checking thoroughly what it does on
on-x86 architectures); so it indeed may produce very suboptimal code.
Still I am somewhat suprised by the code (its even more ugly than I
would expect :-( ) CSE definitly should replace
the increments by a direct usage in addressing (but iirc there is some
code in CSE that sometimes prevents this in order to keep autoinc
working; perhaps this is the case).
I will try to have a look at that.
Zdenek
> Using the rtlopt branch with "-O3 -funroll-loops" produces the following
> assembly which includes loop-invariant address calculations inside
> the loop body: (the loop was unrolled by 8).
>
> _foo:
> lis r3,0xfffe
> stmw r23,-36(r1)
> ori r2,r3,50992
> li r0,1250
> stwux r1,r1,r2
> mtctr r0
> addis r5,r1,0x1
> addi r5,r5,-25504
> L5:
> addi r8,r5,4
> lwz r29,4(r5)
> addi r7,r8,4
> lwz r28,4(r8)
> addi r6,r7,4
> lwz r27,4(r7)
> addi r4,r6,4
> lwz r26,4(r6)
> addi r3,r4,4
> lwz r25,4(r4)
> addi r2,r3,4
> lwz r0,0(r5)
> addi r12,r2,4
> lwz r24,4(r3)
> lwz r23,4(r2)
> ! addis r5,r5,0xffff
> ! addis r9,r8,0xffff
> ! addis r11,r7,0xffff
> ! addis r10,r6,0xffff
> ! addis r8,r4,0xffff
> ! addis r7,r3,0xffff
> ! addis r6,r2,0xffff
> ! addis r4,r12,0xffff
> stw r0,25536(r5)
> stw r29,25536(r9)
> addi r5,r12,4
> stw r28,25536(r11)
> stw r27,25536(r10)
> stw r26,25536(r8)
> stw r25,25536(r7)
> stw r24,25536(r6)
> stw r23,25536(r4)
> bdnz L5
> addi r3,r1,32
> lwz r1,0(r1)
> lmw r23,-36(r1)
> blr
>
> Using "gcc3.4 -O3 -fold-unroll-loops" produces the following assembly. Note
> that here no loop invariant calculations appear in the loop, and also the
> address computation exposes much more ILP than the code above (which
> generates a sequence of 'addi's that depend on one another).
>
> _foo:
> lis r4,0xfffe
> stmw r24,-32(r1)
> ori r2,r4,51008
> stwux r1,r1,r2
> li r2,500
> addi r30,r1,24
> mtctr r2
> addis r3,r30,0x1
> addi r12,r3,-25452
> L48:
> lwz r7,-52(r12)
> lwz r6,-48(r12)
> lwz r5,-44(r12)
> lwz r0,-76(r12)
> lwz r2,-72(r12)
> lwz r9,-68(r12)
> lwz r11,-64(r12)
> lwz r10,-60(r12)
> lwz r8,-56(r12)
> lwz r4,-40(r12)
> lwz r3,-36(r12)
> lwz r29,-32(r12)
> lwz r28,-28(r12)
> lwz r27,-24(r12)
> lwz r26,-20(r12)
> lwz r25,-16(r12)
> lwz r24,-12(r12)
> stw r7,32(r30)
> stw r6,36(r30)
> stw r5,40(r30)
> stw r0,8(r30)
> stw r2,12(r30)
> stw r9,16(r30)
> stw r11,20(r30)
> stw r10,24(r30)
> stw r8,28(r30)
> stw r4,44(r30)
> stw r3,48(r30)
> stw r29,52(r30)
> stw r28,56(r30)
> stw r27,60(r30)
> stw r26,64(r30)
> stw r25,68(r30)
> stw r24,72(r30)
> lwz r7,-8(r12)
> lwz r6,-4(r12)
> lwz r5,0(r12)
> addi r12,r12,80
> stw r7,76(r30)
> stw r6,80(r30)
> stw r5,84(r30)
> addi r30,r30,80
> bdnz L48
> lwz r1,0(r1)
> lmw r24,-32(r1)
>
> Is it indeed the fault of the new unroller?
>
> thanks,
> Yossi
>
More information about the Gcc
mailing list