[Bug c/11327] New: Non-optimal code when using MMX/SSE builtins
kevina at gnu dot org
gcc-bugzilla@gcc.gnu.org
Thu Jun 26 12:44:00 GMT 2003
PLEASE REPLY TO gcc-bugzilla@gcc.gnu.org ONLY, *NOT* gcc-bugs@gcc.gnu.org.
http://gcc.gnu.org/bugzilla/show_bug.cgi?id=11327
Summary: Non-optimal code when using MMX/SSE builtins
Product: gcc
Version: 3.3
Status: UNCONFIRMED
Severity: normal
Priority: P2
Component: c
AssignedTo: unassigned at gcc dot gnu dot org
ReportedBy: kevina at gnu dot org
CC: gcc-bugs at gcc dot gnu dot org
GCC build triplet: i686-pc-linux-gnu
GCC host triplet: i686-pc-linux-gnu
GCC target triplet: i686-pc-linux-gnu
Gcc generates non optimal code when the builtin MMX/SSE functions are used. In
particular it has a tendency to insert unnecessary movq and sometimes
unnecessary memory reads.
For example in the following code:
#include <stddef.h>
typedef int v8qi __attribute__ ((mode(V8QI)));
typedef long long unsigned int ullint;
#define peq __builtin_ia32_pcmpeqb
#define pmin __builtin_ia32_pminub
#define por(a,b) (v8qi)__builtin_ia32_por((ullint)a, (ullint)b)
#define psubs __builtin_ia32_psubusb
void foo(v8qi * a, v8qi * b, v8qi * c, size_t s)
{
size_t i;
v8qi thres;
memset(&thres, 10, 8);
for (i = 0; i != s; ++i)
{
c[i] = peq( pmin( por(psubs(a[i],b[i]), psubs(b[i],a[i])),
thres),
thres);
}
}
When compiled with "-O2 -march=pentium3" Gcc generates:
# %mm2 is the constant thres
...
movq (%ecx,%eax,8), %mm0
psubusb (%edx,%eax,8), %mm0
movq %mm0, %mm1
movq (%edx,%eax,8), %mm0
psubusb (%ecx,%eax,8), %mm0
por %mm0, %mm1
movq %mm1, %mm0
pminub %mm2, %mm0
pcmpeqb %mm2, %mm0
movq %mm0, (%esi,%eax,8)
...
which involves 2 unnecessary memory reads and 1 unnecessary movq. An optimal
version of the above code:
movq (%ecx,%eax,8), %mm0
movq (%edx,%eax,8), %mm1
movq %mm0, %mm3
psubusb %mm1, %mm0
psubusb %mm3, %mm1
por %mm1, %mm0
pminub %mm2, %mm0
pcmpeqb %mm2, %mm0
movq %mm0, (%esi,%eax,8)
This is just a simple example. In more complex code there are more unnecessary
movq. Spelling out exactly what to do for the inner loop:
m1 = a[i];
m2 = b[i];
m3 = m1;
m1 = psubs(m1, m2);
m2 = psubs(m2, m3);
m1 = por(m1,m2);
m1 = pmin(m1, thres);
m1 = peq(m1,thres);
c[i] = m1;
Does not help. It avoids the unnecessary memory reads but adds several
unnecessary movq.
The attached files include the example code and the generated code gcc produces
with a "diff" to my optimal version. The ineffect code is marked with a '-'
while my code is marked with a '+'.
More information about the Gcc-bugs
mailing list