RFC: Prefetch instruction support

Jan Hubicka hubicka@atrey.karlin.mff.cuni.cz
Wed Apr 12 03:09:00 GMT 2000


> In general for Athlon, aggressive unrolling is recommended.
> The ICache is huge (64K). Depending on the loop unrolling
> anywhere from 4x to 16x seems helpful for "typical" code,
> independent from the prefetching issue. YMMV. For example
> in compilers that don't perform scalar replacement unrolling 
> get's you most of the benefits anyhow. The aggressive loop 
> unrolling plays jibes with the goal of having non-overlapping 
> prefetches. I agree you'd probably would want a heuristic
> that prevents the unroller becoming overly aggressive, e.g.
> by considering code size etc. The best solution would be to
> have the unrolling stage benefit from profiling information,
> i.e. feedback directed optimization. Performing loop 
> transformations in order to maximize the number of unit stride 
> accesses is recommended for Athlon (and seems like a sensible 
> policy on just about every CPU short of vector machines with 
> scatter/gather capability).
> 
> I will admit that I have never thought of using prefetch for
> code with large strides, e.g. 8K or some such. To first order

I am plannig to support acceses with large stride soon. The problem
is that analysis are somewhat more complex than I do currently.

My plan is to record all indexes we use to access array with given
striddle and sort them modulo striddle, then identify blocks where
two consetuctive indexes has smaller distance than cache line and
cover them by as many prefetch instructions as neccesary according
the size of block always use largest index in the given block+few
iterations forward.

The heruistics I am using now is much easier to implement, so I planned this to
do later, once the main design of code is stable and I read some relevant
literature.

I would like to support arrays with variable striddle too (ie vertical
traversing of two dimenshional arrays)

Having some support for nested loops traversing two dimensional array can
be sweet too, but I am not sure how to implement this
(otherwise the internal loop will always appear to be small)

> one would think that adding the prefetch could never hurt and 
> would often improve performance. On the other hand couldn't
> it increase the likelihood of conflict misses? I.e., the 
> prefetch could kick data out of the cache which we are currently
> working on. E.g. in a 2-way set associative cache where each
> way is x KByte, prefetching accesses with stride x KByte, several
> units head seems like it could be pretty deadly.

This is interesting point I didn't take into account before.  Perhaps
I can predict such a conflicts (at least inside single array) and
attempt to avoid them somehow.

Honza


More information about the Gcc mailing list