do concurrent in parallel? speed-up?

Tim Prince n8tm@aol.com
Wed Feb 12 14:14:00 GMT 2014


On 2/11/2014 6:22 PM, Tim Prince wrote:
>
> On 2/11/2014 2:31 PM, Damian Rouson wrote:
>> On Tue, 11 Feb 2014 13:15:34 -0500, Tim Prince wrote:
>>
>>> I've turned my masked do concurrents into DO loops with MERGE for 
>>> the benefit of gfortran 4.9. ifort
>>> needs VECTOR ALIGNED directives to get full performance of masked DO 
>>> CONCURRENT, and it's still fully
>>> effective only for single thread vectorization.
>> Could you provide a simple example and give rough estimates of the 
>> resulting speedup with gfortran 4.9 and the CPU employed?
>>
>> Damian
> #if defined __INTEL_COMPILER
> !dir$ vector aligned
>           do concurrent( i= 1:n, a(i) > b(i))
>             a(i)= a(i)-b(i)*d(i)
>             c(i)= a(i)+c(i)
>           enddo
> #else
>           do i= 1,n
>             temp= a(i) > b(i)
>             a(i)= a(i)-merge(b(i)*d(i),0.,temp)
>             c(i)= merge(a(i),0.,temp)+c(i)
>           enddo
> #endif
> That one doesn't vectorize with gfortran. Even without vectorization, 
> the change from do concurrent to merge improves performane by about 
> 30% on several CPU generations. It can be vectorized by gfortran for 
> avx by splitting into individual do loops, but this involves an 
> additional temporary array, so the vectorization would not do much 
> more than double performance, in the absence of loop re-fusion.  Intel 
> Cilk(tm) Plus can achieve excellent performance by that route, but 
> it's very picky about alignment assertions.  As Tobias mentioned, the 
> alignment assertions for avx must match data declarations or compiler 
> data alignment specification (-align array32byte).
> The merges accomplish about what ifort does in optimizing the mask 
> operation by minimizing the required operations.  It prevents use of 
> fused multiply-add; tests show that appears to be the right choice.
> #if defined __INTEL_COMPILER
> !dir$ vector aligned
>           do concurrent( i= 1:n, e(i) >= t)
>             a(i)= a(i)+c(i)*d(i)
>             b(i)= b(i)+c(i)*c(i)
>             enddo
> #else
>           do i= 1,n
>             a(i)= a(i)+merge(c(i)*d(i),0.,e(i) >= t)
>             enddo
>           do i= 1,n
>             b(i)= b(i)+merge(c(i)*c(i),0.,e(i) >= t)
>             enddo
> #endif
> This one gets 85% of potential performance on my avx2 laptop for 
> n==1000.  When built for original corei7, it doesn't vectorize, and 
> the loop split costs some performance then too (repeated common 
> operations).
>
> Both examples are from the classic Levine, Callahan, Dongarra netlib 
> vector benchmark (if you want to see the f66 code).
>
             a(:n)= a(:n)+merge(c(:n)*d(:n),0.,e(:n) >= t)
             b(:n)= b(:n)+merge(c(:n)*c(:n),0.,e(:n) >= t)
also produces about 3x performance of do concurrent with gfortran -O3 
-march=avx2.

-- 
Tim Prince



More information about the Fortran mailing list