Slow MATMUL

N.M. Maclaren nmm1@cam.ac.uk
Sun Dec 12 15:58:00 GMT 2010


On Dec 12 2010, Tobias Burnus wrote:
>
>> One certainly can improve it, but the problems are doing so without
>> introducing system dependencies or reserved names (*gemm).  I suspect
>> that is what you really mean.
>
>Well, gfortran has the option -fexternal-blas, which optionally replaces 
>MATMUL by a call to a user-provided BLAS library, whether it is Netlib's 
>BLAS, the one of ACML, the MKL or ESSL.

I should have checked!  Mea culpa.  That's clearly useful but, as you say,
not a complete solution.

>The question was more what one could do to improve the performance when 
>not using (a vendor) BLAS for MATMUL but the standard libgfortran function.

Unquestionably. The question isn't whether one could but whether one 
should. If I get a moment (ha!), I will try to look at the code, but am not 
very optimistic that ANY code can be both the following:

    Simple, portable and maintainable

    Efficient for all reasonable uses of MATMUL and never dire

A proper test would investigate for a reasonable range of properties, and
I am certain that I could find a flaw in any existing implementation if I
looked hard enough.  It's a heller of a problem, in its full generality.

I looked at an even simpler problem: transposition.  On the 'mainstream'
cached machines, one of the blocked forms was always fastest, but which
one and the block size varied considerably.  On a vector system, hoo, boy!
By FAR the fastest was iterating over the diagonals (i.e. treating the
diagonals as the inner vector).  Those are more-or-less dead, but what
about GPUs?  Embedded systems?  Automatic multi-threaded systems (Niagara,
Larrabee and whatever comes next)?  And how is the speed affected when you
run one copy per core, versus one per system?  And so on ....

And I was considering JUST dense, square matrices.  If I had added the
issue of sparse and rectangular ones (as are generated by sections), then
the whole problem would have got even more complicated.  My suspicion is
that doing a proper test would show up a much less consistent advantage
for NAG and Intel.  Which isn't putting anyone down - it's just the nature
of the problem.


Regards,
Nick Maclaren.



More information about the Fortran mailing list