Efficiency of g77 generated code under egcs-970910
Jeffrey A Law
law@cygnus.com
Wed Sep 17 22:35:00 GMT 1997
In message < 9709141205.AA08272@moene.indiv.nluug.nl >you write:
> Lectoribus Salutem,
>
> Instead of generating yet another set of numbers showing g77 +
> egcs-970910 generated code performing less than that produced by
> g77-0.5.18 (de dato April 1st, 1996), I thought of showing some code
> comparisons:
Thanks!
> This is what we get for the inner loop using g77-0.5.20:
>
> L9:
> fmoved a5@+,fp0
> faddd a0@(-8),fp0
> faddd a3@+,fp0
> faddd a2@+,fp0
> fmoved a0@+,fp1
> fmulx fp3,fp1
> fsubx fp1,fp0
> fmulx fp2,fp0
> fnegx fp0,fp0
> movel d0,a1
> addql #8,d0
> faddd a1@,fp0
> fmoved fp0,a4@+
> dbra d3,L9
This looks like reasonable code, not perfect, but reasonable.
>
> Now g77-0.5.21:
>
> L9:
> fmoved a0@+,fp0
> faddd a3@(d2:l),fp0
> faddd a3@(d4:l),fp0
> faddd a3@(d1:l),fp0
> fmoved a1@(a3:l),fp1
> fmulx fp3,fp1
> fsubx fp1,fp0
> fmulx fp2,fp0
> fnegx fp0,fp0
> movel a6@(16),a4
> faddd a4@(d0:l),fp0
> movel a6@(8),a4
> fmoved fp0,a4@(d3:l)
> addql #8,d2
> addql #8,d0
> addqw #8,a1
> addql #8,d1
> addql #8,d4
> addql #8,d3
> dbra d5,L9
Yick!
What I immediately see is lots of givs that weren't recognized/reduced,
which leads to all the extra address computations in the inner loop.
> Richard Kenner told us (g77-alpha) that this code was removed
> because it actually caused a bug, not just because it's ugly.
> Note: This patch is with respect to gcc-2.7.2.3.
>
> My intuition tells me it is the fold( ... ) calls in the above code
> that cause the difference; unfortunately, one cannot just rerun
> `constant folding' as a separate compiler pass to `prove' this.
Well, let's assume Kenner's right and there's no way to rewrite that change
in a manner that works.
For reference, I get the following on the PA using the latest egcs code:
addl %r28,%r31,%r19
flddx %r19(0,%r25),%fr23
flddx %r22(0,%r25),%fr22
addl %r21,%r4,%r19
fadd,dbl %fr22,%fr23,%fr22
addl %r21,%r31,%r20
flddx %r19(0,%r25),%fr23
addl %r21,%r26,%r19
fadd,dbl %fr22,%fr23,%fr22
flddx %r20(0,%r25),%fr24
flddx %r19(0,%r25),%fr23
fmpyadd,dbl %fr24,%fr27,%fr24,%fr23,%fr22
addl %r21,%r24,%r19
ldo 8(%r28),%r28
flddx %r19(0,%r5),%fr23
fsub,dbl %fr22,%fr24,%fr22
addl %r21,%r23,%r19
ldo 8(%r22),%r22
fmpy,dbl %fr26,%fr22,%fr22
ldo 8(%r21),%r21
fsub,dbl 0,%fr22,%fr22
fadd,dbl %fr22,%fr23,%fr22
addib,>= -1,%r29,L$0009
fstdx %fr22,%r19(0,%r6)
Which, IMHO, sucks.
Looking at things under the debugger, loop.c apparently doesn't recognize
either of these two expressions as invariants (note the "use" signifies
an invariant expression in simplify_giv_expr):
(plus (use (reg) (use (reg))
(plus (use (reg) (const_int))
A quick hack to recognize those expressions as invariants allows loop
to recognize all the givs in the loop and I end up with the following
code for the inner loop on the PA:
fldds,ma 8(0,%r25),%fr22
fldds,ma 8(0,%r23),%fr23
fldds,ma 8(0,%r24),%fr24
fadd,dbl %fr22,%fr23,%fr22
fldds,ma 8(0,%r22),%fr25
fldds,ma 8(0,%r19),%fr23
fmpyadd,dbl %fr23,%fr28,%fr23,%fr24,%fr22
fldds,ma 8(0,%r21),%fr24
fadd,dbl %fr22,%fr25,%fr22
fsub,dbl %fr22,%fr23,%fr22
fmpy,dbl %fr27,%fr22,%fr22
fsub,dbl 0,%fr22,%fr22
fadd,dbl %fr22,%fr24,%fr22
addib,>= -1,%r28,L$0009
fstds,ma %fr22,8(0,%r20)
Which is nearly optimal -- with work we might be able to combine the
fmpy,dbl with one of the fadd or fsub instructions to create fmpysub
and maybe improve the scheduling...
This is an untested patch and probably breaks something; I give it
to y'all in the hopes that someone can make some use out of it:
Index: loop.c
===================================================================
RCS file: /cvs/cvsfiles/egcs/gcc/loop.c,v
retrieving revision 1.11
diff -c -3 -p -r1.11 loop.c
*** loop.c 1997/09/11 17:08:01 1.11
--- loop.c 1997/09/18 05:31:23
*************** simplify_giv_expr (x, benefit)
*** 5368,5383 ****
case CONST_INT:
case USE:
/* Both invariant. Only valid if sum is machine operand.
! First strip off possible USE on first operand. */
if (GET_CODE (arg0) == USE)
arg0 = XEXP (arg0, 0);
tem = 0;
if (CONSTANT_P (arg0) && GET_CODE (arg1) == CONST_INT)
{
tem = plus_constant (arg0, INTVAL (arg1));
if (GET_CODE (tem) != CONST_INT)
tem = gen_rtx (USE, mode, tem);
}
return tem;
--- 5368,5394 ----
case CONST_INT:
case USE:
/* Both invariant. Only valid if sum is machine operand.
! First strip off possible USE on the operands. */
if (GET_CODE (arg0) == USE)
arg0 = XEXP (arg0, 0);
+ if (GET_CODE (arg1) == USE)
+ arg1 = XEXP (arg1, 0);
+
tem = 0;
if (CONSTANT_P (arg0) && GET_CODE (arg1) == CONST_INT)
{
tem = plus_constant (arg0, INTVAL (arg1));
if (GET_CODE (tem) != CONST_INT)
tem = gen_rtx (USE, mode, tem);
+ }
+ else
+ {
+ /* Adding two invariants must result in an invariant,
+ so enclose addition operation inside a USE and
+ return it. I have no clue what might break because
+ of this. */
+ tem = gen_rtx (USE, mode, gen_rtx (PLUS, mode, arg0, arg1));
}
return tem;
More information about the Gcc
mailing list