inadequate multiply-by-const expansion for pentium4
Luchezar Belev
l_belev@yahoo.com
Thu Apr 29 00:48:00 GMT 2004
--- Jim Wilson wrote:
> On Wed, 2004-04-28 at 03:40, Luchezar Belev wrote:
> > up to gcc-3.3.3 the costs of the non-add instructions were doubled, but
> > in gcc-3.4.0 for some reason this was rejected.
>
> It appears the costs were fixed because a problem was noticed with the
> doubled costs. It isn't clear if this was benchmarked though. If you
> can show that the doubled costs benchmark better, then they can be put
> back.
> http://gcc.gnu.org/ml/gcc-patches/2002-10/msg01044.html
>
> > Why to avoid code size expansion when using -O3 or higher?
>
> Because no one else ever noticed or thought of this before. The change
> should be benchmarked though to see if it does give better performance.
>
> > At least this limit could be done in some architecture-dependent maner and
> > tuned more precisely acording to the arch specifics.
>
> Sure, if you can find something that benchmarks well. I suggested using
> both add_cost and shift_cost because we usually get a mixture of shifts
> and adds, and shifts are often more expensive than adds. Using add_cost
> alone might be limiting us to sequences much shorter than 12
> instructions because of shift costs.
> --
> Jim Wilson, GNU Tools Support, http://www.SpecifixInc.com
Hi, I would try to make changes and patches for gcc myself
but I'm not a big deal familiar with the gcc internals,
so I'm afraid to take up this task.
Originally I meant instead to call to this issue the attantion
of someone else who knows well how exactly these things work in gcc.
BTW here's a a table with instruction latencies for pentium4
and the stupid program I used to measure them, if someone is interested
------------------------------------------------------------
#include <stdio.h>
/*
* The latencies are in cycles.
* Only register-register variants of these instructions were tested
* (no memory operands);
* For the non-marked with `~' numbers, I get quite consistent
* results (often close to the showed numbers with 2 or 3 digit after
* the decimal point, but always greater or equal to them), so
* their accuracy seems quite probable to me;
* For those marked with `~', the exact value is probably somewhere
* around the showed number but the test results vary in such a way,
* that I can't be sure;
* For the cmp instruction I get a latency of around 0.35 (less than 0.5)
* cycles which is a mistery for me.
*
* ?0.35: cmp
* 0.5: mov, add, sub, and, or, xor, neg, not, nop, test, cwde
* ~0.5: sahf
* 1.0: lea, inc, dec, set<cc>, xchg, jmp, predicted_j<cc>, cwd, cdq
* 2.0: predicted_j[e]cxz
* 4.0: shl_{im,1}, shr_{im,1}, sar_{im,1}, rol_{im,1}, ror_{im,1}, rcl_1, rcr_1, bsf, bsr, lahf
* 5.0: bt
* 6.0: btr, bts, btc, shl_cl, shr_cl, sar_cl, rol_cl, ror_cl, cmov<cc>
* 7.0: adc, sbb, bswap
* ~8: mispredicted_j[e]cxz
* 9.0: clc, stc, cmc
* 10.0 emms
* 12.0: shld, shrd
* 14.0: imul_im
* ~18: rcl_{im,cl}, rcr_{im,cl}
* ~21: mispredicted_j<cc>
* 36.5: lfence, sfence
* 46.0: std
* 50.0: cld
* ~80: rdtsc
* ~100: mfence
*/
/*
* 1.0: movaps, movaps_frommem
* 2.0: movups_frommem, mofaps_tomem
* 8.0: movups_tomem
*/
//#define TEST_INSN "imull $13, %%eax ; "
//#define CLOB_REGS "eax",
//#define TEST_INSN "inc %%eax ;"
//#define CLOB_REGS "eax",
#define REPEAT 64*32
#define rdtsc() \
({ \
union { \
struct { \
unsigned l, h; \
} p; \
unsigned long long f; \
} t; \
asm volatile ("rdtsc" : "=a" (t.p.l), "=d" (t.p.h)); \
t.f; \
})
int main()
{
unsigned long long t1, t2;
int k;
t1 = rdtsc();
asm volatile (
".balign 16\n\t"
"1:\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
"subl $1, %%ebx\n\t"
"jnz 1b"
: "=b" (k) : "0" (REPEAT) : CLOB_REGS "cc");
t2 = rdtsc();
printf("cycles per op: %f\n", (double)(t2 - t1)/(REPEAT * 64));
return 0;
}
__________________________________
Do you Yahoo!?
Win a $20,000 Career Makeover at Yahoo! HotJobs
http://hotjobs.sweepstakes.yahoo.com/careermakeover
More information about the Gcc
mailing list