Performance regression with Eigen for non-AVX2 CPUs

Patrik Huber patrikhuber@gmail.com
Wed Feb 7 15:03:00 GMT 2018


Hello,

I noticed today what may look like quite a large performance regression
with Eigen (3.3.4) matrix multiplication. It only seems to occur on
non-AVX2 code paths, meaning that if I compile with -march=native on my
core-i7 with AVX2, then it's blazingly fast on both g++ versions, but not
on an older core-i5 with only AVX, or if I use -march=core2.

Here are some example timings, but it applies to all matrix sizes that the
benchmark script tests (see end of the message for the code):

g++-5 gemm_test.cpp -std=c++17 -I 3rdparty/eigen/ -march=core2 -O3 -o
gcc5_gemm_test

1124 1215 1465
elapsed_ms: 1970
--------
1730 1235 1758
elapsed_ms: 3505

g++-7 gemm_test.cpp -std=c++17 -I 3rdparty/eigen/ -march=core2 -O3
-march=core2 -o gcc7_gemm_test

1124 1215 1465
elapsed_ms: 2998
--------
1730 1235 1758
elapsed_ms: 4628

It's even worse if I test this on a i5-3550, which has AVX, but not AVX2:

g++-5 gemm_test.cpp -std=c++17 -I 3rdparty/eigen/ -march=native -O3 -o
gcc5_gemm_test
1124 1215 1465
elapsed_ms: 941
--------
1730 1235 1758
elapsed_ms: 1780


g++-7 gemm_test.cpp -std=c++17 -I 3rdparty/eigen/ -march=native -O3 -o
gcc7_gemm_test

1124 1215 1465
elapsed_ms: 1988
--------
1730 1235 1758
elapsed_ms: 3740

I tried the same with -O2 and it gave the same results. That's a drop to
nearly half the speed in matrix multiplication on AVX CPUs. Or maybe I've
done something wrong. :-) I realise the benchmark might be a bit crude
(better use Google Benchmark or something like that...) But the results I'm
getting are pretty consistent on various CPUs, compilers, and with various
flags.

Best wishes,

Patrik

===
// gemm_test.cpp
#include <array>
#include <chrono>
#include <iostream>
#include <random>
#include <Eigen/Dense>

using RowMajorMatrixXf = Eigen::Matrix<float, Eigen::Dynamic,
Eigen::Dynamic, Eigen::RowMajor>;
using ColMajorMatrixXf = Eigen::Matrix<float, Eigen::Dynamic,
Eigen::Dynamic, Eigen::ColMajor>;

template <typename Mat>
void run_test(const std::string& name, int s1, int s2, int s3)
{
    using namespace std::chrono;
    float checksum = 0.0f; // to prevent compiler from optimizing
everything away
    const auto start_time_ns =
high_resolution_clock::now().time_since_epoch().count();
    for (size_t i = 0; i < 10; ++i)
    {
        Mat a_rm(s1, s2);
        Mat b_rm(s2, s3);
        const auto c_rm = a_rm * b_rm;
        checksum += c_rm(0, 0);
    }
    const auto end_time_ns =
high_resolution_clock::now().time_since_epoch().count();
    const auto elapsed_ms = (end_time_ns - start_time_ns) / 1000000;
    std::cout << name << " (checksum: " << checksum << ") elapsed_ms: " <<
elapsed_ms << std::endl;
}
int main()
{
    //std::random_device rd;
    //std::mt19937 gen(0);
    //std::uniform_int_distribution<> dis(1, 2048);
    std::vector<int> vals = { 1124, 1215, 1465, 1730, 1235, 1758, 1116,
1736, 868, 1278, 1323, 788 };
    for (std::size_t i = 0; i < 12; ++i)
    {
        int s1 = vals[i++];//dis(gen);
        int s2 = vals[i++];//dis(gen);
        int s3 = vals[i];//dis(gen);
        std::cout << s1 << " " << s2 << " " << s3 << std::endl;
        run_test<ColMajorMatrixXf>("col major", s1, s2, s3);
        run_test<RowMajorMatrixXf>("row major", s1, s2, s3);
        std::cout << "--------" << std::endl;
    }
    return 0;
}
===

-- 
Dr. Patrik Huber
Centre for Vision, Speech and Signal Processing
University of Surrey
Guildford, Surrey GU2 7XH
United Kingdom

Web: www.patrikhuber.ch
Mobile: +44 (0)7482 633 934



More information about the Gcc-regression mailing list