This is the mail archive of the
gcc-bugs@gcc.gnu.org
mailing list for the GCC project.
Strange behavior of gcc 2.95.2
- To: gcc-bugs at gcc dot gnu dot org
- Subject: Strange behavior of gcc 2.95.2
- From: Regis Duchesne <hpreg at vmware dot com>
- Date: 13 Mar 2000 17:24:59 -0800
- Cc: mts at vmware dot com
Hi guys,
I used to compile my kernel code with gcc 2.95.2. Recently, a co-worker
compiled it gcc 2.7.2.3. Each time he was running my kernel code, his
machine was resetting.
I look further and narrowed the problem to the following sha1 code:
----------- Cut here ---------
/* gcc -g -Wall -O6 -c testgcc.c */
extern inline void *
memcpy(void *destPtr, const void *srcPtr, unsigned int count)
{
char *tmp = destPtr;
const char *src = srcPtr;
while (count--)
*tmp++ = *src++;
return destPtr;
}
#define rol(value, bits) (((value) << (bits)) | ((value) >> (32 - (bits))))
/* blk0() and blk() perform the initial expand. */
/* I got the idea of expanding during the round function from SSLeay */
#define blk0(i) (block->l[i] = (rol(block->l[i],24)&0xFF00FF00) \
|(rol(block->l[i],8)&0x00FF00FF))
#define blk(i) (block->l[i&15] = rol(block->l[(i+13)&15]^block->l[(i+8)&15] \
^block->l[(i+2)&15]^block->l[i&15],1))
/* (R0+R1), R2, R3, R4 are the different operations used in SHA1 */
#define R0(v,w,x,y,z,i) z+=((w&(x^y))^y)+blk0(i)+0x5A827999+rol(v,5);w=rol(w,30);
#define R1(v,w,x,y,z,i) z+=((w&(x^y))^y)+blk(i)+0x5A827999+rol(v,5);w=rol(w,30);
#define R2(v,w,x,y,z,i) z+=(w^x^y)+blk(i)+0x6ED9EBA1+rol(v,5);w=rol(w,30);
#define R3(v,w,x,y,z,i) z+=(((w|x)&y)|(w&x))+blk(i)+0x8F1BBCDC+rol(v,5);w=rol(w,30);
#define R4(v,w,x,y,z,i) z+=(w^x^y)+blk(i)+0xCA62C1D6+rol(v,5);w=rol(w,30);
/* Hash a single 512-bit block. This is the core of the algorithm. */
static void SHA1Transform(unsigned long state[5], unsigned char buffer[64])
{
typedef union {
unsigned char c[64];
unsigned long l[16];
} CHAR64LONG16;
unsigned long a, b, c, d, e;
CHAR64LONG16* block;
#if 0
static unsigned char workspace[64];
/* gcc 2.7.2.3: stack frame is
0x14 bytes if the memcpy comes last
0x514 bytes if the memcpy comes first */
/* gcc 2.95.2: stack frame is
0x15c bytes if the memcpy comes last
0x11c bytes if the memcpy comes first */
block = (CHAR64LONG16*)workspace;
memcpy(workspace, buffer, 64);
#else
/* gcc 2.7.2.3 stack frame is 0x14 bytes */
/* gcc 2.95.2 stack frame is 0x15c bytes */
block = (CHAR64LONG16*)buffer;
#endif
/* Copy context->state[] to working vars */
a = state[0];
b = state[1];
c = state[2];
d = state[3];
e = state[4];
/* 4 rounds of 20 operations each. Loop unrolled. */
R0(a,b,c,d,e, 0); R0(e,a,b,c,d, 1); R0(d,e,a,b,c, 2); R0(c,d,e,a,b, 3);
R0(b,c,d,e,a, 4); R0(a,b,c,d,e, 5); R0(e,a,b,c,d, 6); R0(d,e,a,b,c, 7);
R0(c,d,e,a,b, 8); R0(b,c,d,e,a, 9); R0(a,b,c,d,e,10); R0(e,a,b,c,d,11);
R0(d,e,a,b,c,12); R0(c,d,e,a,b,13); R0(b,c,d,e,a,14); R0(a,b,c,d,e,15);
R1(e,a,b,c,d,16); R1(d,e,a,b,c,17); R1(c,d,e,a,b,18); R1(b,c,d,e,a,19);
R2(a,b,c,d,e,20); R2(e,a,b,c,d,21); R2(d,e,a,b,c,22); R2(c,d,e,a,b,23);
R2(b,c,d,e,a,24); R2(a,b,c,d,e,25); R2(e,a,b,c,d,26); R2(d,e,a,b,c,27);
R2(c,d,e,a,b,28); R2(b,c,d,e,a,29); R2(a,b,c,d,e,30); R2(e,a,b,c,d,31);
R2(d,e,a,b,c,32); R2(c,d,e,a,b,33); R2(b,c,d,e,a,34); R2(a,b,c,d,e,35);
R2(e,a,b,c,d,36); R2(d,e,a,b,c,37); R2(c,d,e,a,b,38); R2(b,c,d,e,a,39);
R3(a,b,c,d,e,40); R3(e,a,b,c,d,41); R3(d,e,a,b,c,42); R3(c,d,e,a,b,43);
R3(b,c,d,e,a,44); R3(a,b,c,d,e,45); R3(e,a,b,c,d,46); R3(d,e,a,b,c,47);
R3(c,d,e,a,b,48); R3(b,c,d,e,a,49); R3(a,b,c,d,e,50); R3(e,a,b,c,d,51);
R3(d,e,a,b,c,52); R3(c,d,e,a,b,53); R3(b,c,d,e,a,54); R3(a,b,c,d,e,55);
R3(e,a,b,c,d,56); R3(d,e,a,b,c,57); R3(c,d,e,a,b,58); R3(b,c,d,e,a,59);
R4(a,b,c,d,e,60); R4(e,a,b,c,d,61); R4(d,e,a,b,c,62); R4(c,d,e,a,b,63);
R4(b,c,d,e,a,64); R4(a,b,c,d,e,65); R4(e,a,b,c,d,66); R4(d,e,a,b,c,67);
R4(c,d,e,a,b,68); R4(b,c,d,e,a,69); R4(a,b,c,d,e,70); R4(e,a,b,c,d,71);
R4(d,e,a,b,c,72); R4(c,d,e,a,b,73); R4(b,c,d,e,a,74); R4(a,b,c,d,e,75);
R4(e,a,b,c,d,76); R4(d,e,a,b,c,77); R4(c,d,e,a,b,78); R4(b,c,d,e,a,79);
/* Add the working vars back into context.state[] */
state[0] += a;
state[1] += b;
state[2] += c;
state[3] += d;
state[4] += e;
/* Wipe variables */
a = b = c = d = e = 0;
}
----------- Cut here ---------
The problem is a stack smash upon entry of the SHA1Transform stack
frame (which leads to a triple fault). It seems that under some
circumstances, gcc 2.7.2.3 and gcc 2.95.2 allocate a lot of memory on the stack, while only a fraction of it is actually necessary:
To find the size of the stack frame, I used gdb (disas SHA1Transform)
on the .o generated with gcc -g -Wall -O6 -c testgcc.c
Summary:
compiler | call to memcpy | memcpy comes first | stack size used
----------------------------------------------------------------
2.7.2.3 N N/A 0x014
2.95.2 N N/A 0x15c
2.7.2.3 Y N 0x514
2.95.2 Y N 0x15c
2.7.2.3 Y Y 0x014
2.95.2 Y Y 0x11c
Some numbers are outrageously high (0x514 actually caused the stack smashed
that allowed us to notice the bug) with gcc 2.7.2.3, but they are
still quite high with gcc 2.95.2.
Independantly of the memcpy() call or of its ordering, I expected a
stack size that is about the size of all the local variables,
i.e. maximum 0x18. Maybe I'm missing something here, but it appears to
me that stack sizes that are close to 16 times larger than expected
are bugs.
Don't hesitate to send me your comments or requests of additional
info on this.
Thanks in advance,
Keep up the good gcc work,
--
Regis "HPReg" Duchesne - Member of Technical Staff - VMWare, Inc.
www http://www.VMware.com/
(O o) I use Linux (1135 KB/s over 10Mb/s ethernet)
--.oOO--(_)--OOo.----------------------------------------------------
If cryptography is outlawed, only outlaws will have cryptography