[Bug target/82136] x86: -mavx256-split-unaligned-load should try to fold other shuffles into the load/vinsertf128
peter at cordes dot ca
gcc-bugzilla@gcc.gnu.org
Fri Sep 8 02:08:00 GMT 2017
https://gcc.gnu.org/bugzilla/show_bug.cgi?id=82136
--- Comment #1 from Peter Cordes <peter at cordes dot ca> ---
Whoops, the compiler-explorer link had aligned=1. This one produces the asm I
showed in the original report: https://godbolt.org/g/WsZ5S9
See bug 82137 for a much more efficient vectorization strategy gcc should use
instead, with just in-lane shuffle + blend and some duplicated work.
More information about the Gcc-bugs
mailing list