[PATCH v3] RISC-V: add multiarch RVV support for memcpy using FMV IFUNC
enh
enh@google.com
Wed Jan 22 14:30:47 GMT 2025
On Wed, Jan 22, 2025 at 8:25 AM Adhemerval Zanella Netto
<adhemerval.zanella@linaro.org> wrote:
>
>
>
> On 22/01/25 09:53, Adhemerval Zanella Netto wrote:
> >
> >
> > On 21/01/25 06:45, daichengrong wrote:
> >>
> >> 在 2025/1/20 20:55:09, Adhemerval Zanella Netto 写道:
> >>> On 20/01/25 00:26, daichengrong@iscas.ac.cn wrote:
> >>>> From: daichengrong <daichengrong@iscas.ac.cn>
> >>>>
> >>>> This patch introduces vector support for memcpy with IFUNC.
> >>>> The implementation select the RVV optimized memcpy version via hwprobe on kernel with RVV enabled.
> >>>>
> >>>> Changes in v3:
> >>>> Remove unnecessary whitespace
> >>>>
> >>>> Changes in v2:
> >>>> delete size-0 branch
> >>>>
> >>>> Signed-off-by: daichengrong <daichengrong@iscas.ac.cn>
> >>> For performance improvements the usual way is to also provide some benchmark
> >>> numbers on real hardware (either using our synthetic benchtests or some other
> >>> benchmark that stresses the affected routines).
> >> Thanks! It runs bench-memcpy at K230.
> >>> Also, I take that RISCV_HWPROBE_IMA_V implies in RISCV_HWPROBE_MISALIGNED_FAST
> >>> right?
> >>
> >> SPEC says :"Support for misaligned vector memory accesses is independent of an implementation’s support for misaligned scalar memory accesses. "
> >
> > This kernel patchset [1] hints that RVV implementation suffers the same issue of
> > scalar wrt memory access. The RISCV_HWPROBE_MISALIGNED_VECTOR_SLOW is troublesome
> > because it means that misaligned access might be trapped by the kernel (which
> > might render the implementation really slow); but RISCV_HWPROBE_MISALIGNED_VECTOR_UNSUPPORTED
> > is worrisome because it means this implementation will be just wrong in this
> > kind of hardware.
> >
> > So maybe it would be better to either adapt the implementation to only use aligned
> > memory access (with head and tail handling); or wait this kernel patchset to land
> > and enable it only for RISCV_HWPROBE_MISALIGNED_VECTOR_FAST.
> >
> > And I am pressing on performance numbers on real hardware because on different
> > hardware there is usually some conditions on when vector instruction do show
> > better performance on string routines (either alignment or size). Even for
> > x86_64 ERMS on very recent chips requires tuning to avoid performance pitfalls
> > [2] [3]. Maybe RVV requires that this kind of memory access act like AArch64
>
> This is confunsing, I meant 'Maybe RVV implementations do not require such tuning
> and work like AArch64 MOPS, ...'.
no, anything is possible... the finer-grained distinctions came about
_after_ we saw hardware with fast unaligned scalar but slow unaligned
vector.
> > MOPS, where it is up to hardware to figure out; although the kernel patchset
> > hinted me that it might not be the case for all implementations.
> >
> > Keep in mind that you are proposing a generic implementation, meaning once
> > enabled it would be used for all RISCV chips and not only for K230. It means
> > that we need to be careful to proper enable and make it future proof.
> >
> > [1] https://patchwork.kernel.org/project/linux-riscv/cover/20241017-jesse_unaligned_vector-v10-0-5b33500160f8@rivosinc.com/
> > [2] https://sourceware.org/bugzilla/show_bug.cgi?id=30994
> > [3] https://sourceware.org/bugzilla/show_bug.cgi?id=32475
>
More information about the Libc-alpha
mailing list