Slow MATMUL
N.M. Maclaren
nmm1@cam.ac.uk
Sun Dec 12 15:58:00 GMT 2010
On Dec 12 2010, Tobias Burnus wrote:
>
>> One certainly can improve it, but the problems are doing so without
>> introducing system dependencies or reserved names (*gemm). I suspect
>> that is what you really mean.
>
>Well, gfortran has the option -fexternal-blas, which optionally replaces
>MATMUL by a call to a user-provided BLAS library, whether it is Netlib's
>BLAS, the one of ACML, the MKL or ESSL.
I should have checked! Mea culpa. That's clearly useful but, as you say,
not a complete solution.
>The question was more what one could do to improve the performance when
>not using (a vendor) BLAS for MATMUL but the standard libgfortran function.
Unquestionably. The question isn't whether one could but whether one
should. If I get a moment (ha!), I will try to look at the code, but am not
very optimistic that ANY code can be both the following:
Simple, portable and maintainable
Efficient for all reasonable uses of MATMUL and never dire
A proper test would investigate for a reasonable range of properties, and
I am certain that I could find a flaw in any existing implementation if I
looked hard enough. It's a heller of a problem, in its full generality.
I looked at an even simpler problem: transposition. On the 'mainstream'
cached machines, one of the blocked forms was always fastest, but which
one and the block size varied considerably. On a vector system, hoo, boy!
By FAR the fastest was iterating over the diagonals (i.e. treating the
diagonals as the inner vector). Those are more-or-less dead, but what
about GPUs? Embedded systems? Automatic multi-threaded systems (Niagara,
Larrabee and whatever comes next)? And how is the speed affected when you
run one copy per core, versus one per system? And so on ....
And I was considering JUST dense, square matrices. If I had added the
issue of sparse and rectangular ones (as are generated by sections), then
the whole problem would have got even more complicated. My suspicion is
that doing a proper test would show up a much less consistent advantage
for NAG and Intel. Which isn't putting anyone down - it's just the nature
of the problem.
Regards,
Nick Maclaren.
More information about the Fortran
mailing list