Transformations to increase parallelism (peepholes?)

Jan Hoogerbrugge hoogerbrugge@hotmail.com
Wed Jul 23 16:03:00 GMT 2003


> > Hi,
> >
> > Are there optimizations in gcc that increase instruction level 
>parallelism?
> > For example:
> >
> > void foo(int *p, int a)
> > {
> >         a += 2;
> >         p[0] = a;
> >         a += 2;
> >         p[1] = a;
> > }
> >
> > compiles to
> >
> >         add a, 2 -> tmp1
> >         store p[0], tmp1
> >         add tmp1, 2 -> tmp2
> >         store p[1], tmp2
> >
> > However, a more parallel translation would be:
> >
> >         add a, 2 -> tmp1
> >         store p[0], tmp1
> >         add a, 4 -> tmp2
> >         store p[1], tmp2
> >
> > In this case the two adds and the two stores can be executed in 
>parallel.
> >
> > Jan
>
>No. I've mentioned similar problems before on this list, though.
>
>GCC really needs a pass to preprocess the instructions before sched1 to
>give the scheduler more scheduling freedom, especially on many-issue
>processors.


Would it be possible to do (some of) these transformations by means of 
peepholes in the .md file? If so, could somebody tell me how a peephole 
should look like for

   reg1 = reg2 + const1
   reg3 = reg1 + const2

to

   reg1 = reg2 + const1
   reg3 = reg2 + (const1 + const2)

where const1 + const2 has to be within certain bounds. It tried to write a 
peephole for this but without success. Who helps?

Cheers,
Jan

_________________________________________________________________
Help STOP SPAM with the new MSN 8 and get 2 months FREE*  
http://join.msn.com/?page=features/junkmail



More information about the Gcc mailing list