inadequate multiply-by-const expansion for pentium4

Luchezar Belev l_belev@yahoo.com
Thu Apr 29 00:48:00 GMT 2004


--- Jim Wilson wrote:
> On Wed, 2004-04-28 at 03:40, Luchezar Belev wrote:
> > up to gcc-3.3.3 the costs of the non-add instructions were doubled, but
> > in gcc-3.4.0 for some reason this was rejected.
> 
> It appears the costs were fixed because a problem was noticed with the
> doubled costs.  It isn't clear if this was benchmarked though.  If you
> can show that the doubled costs benchmark better, then they can be put
> back.
>   http://gcc.gnu.org/ml/gcc-patches/2002-10/msg01044.html
> 
> > Why to avoid code size expansion when using -O3 or higher?
> 
> Because no one else ever noticed or thought of this before.  The change
> should be benchmarked though to see if it does give better performance.
> 
> > At least this limit could be done in some architecture-dependent maner and
> > tuned more precisely acording to the arch specifics.
> 
> Sure, if you can find something that benchmarks well.  I suggested using
> both add_cost and shift_cost because we usually get a mixture of shifts
> and adds, and shifts are often more expensive than adds.  Using add_cost
> alone might be limiting us to sequences much shorter than 12
> instructions because of shift costs.
> -- 
> Jim Wilson, GNU Tools Support, http://www.SpecifixInc.com


Hi, I would try to make changes and patches for gcc myself
but I'm not a big deal familiar with the gcc internals,
so I'm afraid to take up this task.
Originally I meant instead to call to this issue the attantion
of someone else who knows well how exactly these things work in gcc.


BTW here's a a table with instruction latencies for pentium4
and the stupid program I used to measure them, if someone is interested


------------------------------------------------------------

#include <stdio.h>


/*
 * The latencies are in cycles.
 * Only register-register variants of these instructions were tested
 * (no memory operands);
 * For the non-marked with `~' numbers, I get quite consistent
 * results (often close to the showed numbers with 2 or 3 digit after
 * the decimal point, but always greater or equal to them), so
 * their accuracy seems quite probable to me;
 * For those marked with `~', the exact value is probably somewhere
 * around the showed number but the test results vary in such a way,
 * that I can't be sure;
 * For the cmp instruction I get a latency of around 0.35 (less than 0.5)
 * cycles which is a mistery for me.
 *
 * ?0.35: cmp
 * 0.5:   mov, add, sub, and, or, xor, neg, not, nop, test, cwde
 * ~0.5:  sahf
 * 1.0:   lea, inc, dec, set<cc>, xchg, jmp, predicted_j<cc>, cwd, cdq
 * 2.0:   predicted_j[e]cxz
 * 4.0:   shl_{im,1}, shr_{im,1}, sar_{im,1}, rol_{im,1}, ror_{im,1}, rcl_1, rcr_1, bsf, bsr, lahf
 * 5.0:   bt
 * 6.0:   btr, bts, btc, shl_cl, shr_cl, sar_cl, rol_cl, ror_cl, cmov<cc>
 * 7.0:   adc, sbb, bswap
 * ~8:    mispredicted_j[e]cxz
 * 9.0:   clc, stc, cmc
 * 10.0   emms
 * 12.0:  shld, shrd
 * 14.0:  imul_im
 * ~18:   rcl_{im,cl}, rcr_{im,cl}
 * ~21:   mispredicted_j<cc>
 * 36.5:  lfence, sfence
 * 46.0:  std
 * 50.0:  cld
 * ~80:   rdtsc
 * ~100:  mfence
 */

/*
 * 1.0:   movaps, movaps_frommem
 * 2.0:   movups_frommem, mofaps_tomem
 * 8.0:   movups_tomem
 */




//#define TEST_INSN "imull $13, %%eax ; "
//#define CLOB_REGS "eax",

//#define TEST_INSN "inc %%eax ;"
//#define CLOB_REGS "eax",


#define REPEAT 64*32




#define rdtsc() \
({ \
	union { \
		struct { \
			unsigned l, h; \
		} p; \
		unsigned long long f; \
	} t; \
	asm volatile ("rdtsc" : "=a" (t.p.l), "=d" (t.p.h)); \
	t.f; \
})



int main()
{
	unsigned long long t1, t2;
	int k;
	
	t1 = rdtsc();
	asm volatile (

		".balign 16\n\t"
		"1:\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"

		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"

		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"

		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"
		TEST_INSN TEST_INSN TEST_INSN TEST_INSN "\n\t"

		"subl	$1, %%ebx\n\t"
		"jnz	1b"
		: "=b" (k) : "0" (REPEAT) : CLOB_REGS "cc");
	t2 = rdtsc();

	printf("cycles per op: %f\n", (double)(t2 - t1)/(REPEAT * 64));

	return 0;
}





 


	
		
__________________________________
Do you Yahoo!?
Win a $20,000 Career Makeover at Yahoo! HotJobs  
http://hotjobs.sweepstakes.yahoo.com/careermakeover 



More information about the Gcc mailing list