(a+b)+c should be replaced by a+(b+c)

Robert Dewar dewar@gnat.com
Thu Mar 25 14:45:00 GMT 2004


Joost VandeVondele wrote:

> good compilers (e.g. xlf90) will (at -O4) do higher order transforms of
> the loop to introduce blocking, independent FMAs, ... that makes this
> little piece of code about 100 times faster at O4 than O2 (what about
> LNO/SSA?). This can only be done if you allow (a+b)+c -> a+(b+c). It is
> basically what any optimized blas routine will do. Matrix multiply is a
> trivial example, if you want blas performance, call blas. There are many
> other kernels like this in e.g. scientific code that are not blas. You
> can't expect a scientist to hand unroll and block any kernel to the
> appropriate depth for any machine. There need to be a compiler option to
> do this. This can only be done if you allow (a+b)+c -> a+(b+c).

Can you really deduce this freedom from later versions of the Fortran
standard?



More information about the Gcc mailing list