Intrinsics for absolute minimum and maximum, and table lookup #324

momchil-velikov · 2024-06-13T16:10:13Z

name: Pull request
about: Technical issues, document format problems, bugs in scripts or feature proposal.

Thank you for submitting a pull request!

If this PR is about a bugfix:

Please use the bugfix label and make sure to go through the checklist below.

If this PR is about a proposal:

We are looking forward to evaluate your proposal, and if possible to
make it part of the Arm C Language Extension (ACLE) specifications.

We would like to encourage you reading through the contribution
guidelines, in particular the section on submitting
a proposal.

Please use the proposal label.

As for any pull request, please make sure to go through the below
checklist.

Checklist: (mark with X those which apply)

If an issue reporting the bug exists, I have mentioned it in the
PR (do not bother creating the issue if all you want to do is
fixing the bug yourself).
I have added/updated the SPDX-FileCopyrightText lines on top
of any file I have edited. Format is SPDX-FileCopyrightText: Copyright {year} {entity or name} <{contact informations}>
(Please update existing copyright lines if applicable. You can
specify year ranges with hyphen , as in 2017-2019, and use
commas to separate gaps, as in 2018-2020, 2022).
I have updated the Copyright section of the sources of the
specification I have edited (this will show up in the text
rendered in the PDF and other output format supported). The
format is the same described in the previous item.
I have run the CI scripts (if applicable, as they might be
tricky to set up on non-*nix machines). The sequence can be
found in the contribution
guidelines. Don't
worry if you cannot run these scripts on your machine, your
patch will be automatically checked in the Actions of the pull
request.
I have added an item that describes the changes I have
introduced in this PR in the section Changes for next
release of the section Change Control/Document history
of the document. Create Changes for next release if it does
not exist. Notice that changes that are not modifying the
content and rendering of the specifications (both HTML and PDF)
do not need to be listed.
When modifying content and/or its rendering, I have checked the
correctness of the result in the PDF output (please refer to the
instructions on how to build the PDFs
locally).
The variable draftversion is set to true in the YAML header
of the sources of the specifications I have modified.
Please DO NOT add my GitHub profile to the list of contributors
in the README page of the project.

neon_intrinsics/advsimd.md

main/acle.md

This patch adds these intrinsics: // Variants are also available for: // [_s8], [_u16], [_s16], [_u32], [_s32], [_u64], [_s64] // [_bf16], [_f16], [_f32], [_f64] void svwrite_lane_zt[_u8](uint64_t zt0, svuint8_t zt, uint64_t idx) __arm_streaming __arm_inout("zt0"); void svwrite_zt[_u8](uint64_t zt0, svuint8_t zt) __arm_streaming __arm_inout("zt0"); according to PR#324[1] [1]ARM-software/acle#324

momchil-velikov · 2024-07-04T17:14:38Z

main/acle.md

+Lookup table read with 4-bit indexes and 8-bit elements.
+``` c
+  // Variants are also available for: _s8
+  svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0");


No, the zn contains indices, it has fixed type, cannot be used for overloading. E.g. the variant for s8 would be

svint8x4_t svluti4_zt_s8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0");

This patch adds these intrinsics: // Variants are also available for: _s8 svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0"); according to PR#324[1] [1]ARM-software/acle#324

tools/intrinsic_db/advsimd.csv

andrewcarlotti

This patch is missing the SME2 FAMINMAX intrinsics

tools/intrinsic_db/advsimd.csv

This patch implements the intrinsics of the form floatNxM_t vamin[q]_fN(floatNxM_t vn, floatNxM_t vm); floatNxM_t vamax[q]_fN(floatNxM_t vn, floatNxM_t vm); as defined in ARM-software/acle#324 Co-authored-by: Hassnaa Hamdi <[email protected]>

This patch implements the following intrinsics: * Floating-point absolute maximum (predicated) svfloat16_t svamax[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_z(svbool_t, svfloat16_t, float16_t); * Floating-point absolute minimum (predicated) svfloat16_t svmin[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_z(svbool_t, svfloat16_t, float16_t); All the intrinsics have also variants for `f32` and `f64`, and have the `__arm_streaming` attribute. (cf. ARM-software/acle#324)

This patch implements these intrinsics: ``` c // Variants are also available for: // [_f32_x2], [_f64_x2], // [_f16_x4], [_f32_x4], [_f64_x4] svfloat16x2_t svamax[_f16_x2](svfloat16x2 zd, svfloat16x2_t zm) __arm_streaming; svfloat16x2_t svamin[_f16_x2](svfloat16x2 zd, svfloat16x2_t zm) __arm_streaming; ``` (cf. ARM-software/acle#324) Co-authored-by: Caroline Concatto <[email protected]>

main/acle.md

tools/intrinsic_db/advsimd.csv

rsandifo-arm · 2024-07-31T13:40:13Z

LGTM.

vhscampos · 2024-07-31T14:17:00Z

Thanks for the PR. I can merge it once the conflicts have been resolved.

vhscampos · 2024-08-05T08:20:23Z

There's one little issue in one Copyright header. Please check the build log to spot the problem.

…ntly

This patch implements the following intrinsics: * Floating-point absolute maximum (predicated) svfloat16_t svamax[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_z(svbool_t, svfloat16_t, float16_t); * Floating-point absolute minimum (predicated) svfloat16_t svmin[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_z(svbool_t, svfloat16_t, float16_t); All the intrinsics have also variants for `f32` and `f64`, and have the `__arm_streaming` attribute. (cf. ARM-software/acle#324)

This patch implements these intrinsics: ``` c // Variants are also available for: // [_f32_x2], [_f64_x2], // [_f16_x4], [_f32_x4], [_f64_x4] svfloat16x2_t svamax[_f16_x2](svfloat16x2 zd, svfloat16x2_t zm) __arm_streaming; svfloat16x2_t svamin[_f16_x2](svfloat16x2 zd, svfloat16x2_t zm) __arm_streaming; ``` (cf. ARM-software/acle#324) Co-authored-by: Caroline Concatto <[email protected]>

This patch implements the intrinsics of the form floatNxM_t vamin[q]_fN(floatNxM_t vn, floatNxM_t vm); floatNxM_t vamax[q]_fN(floatNxM_t vn, floatNxM_t vm); as defined in ARM-software/acle#324 Co-authored-by: Hassnaa Hamdi <[email protected]>

This patch implements the intrinsics of the form floatNxM_t vamin[q]_fN(floatNxM_t vn, floatNxM_t vm); floatNxM_t vamax[q]_fN(floatNxM_t vn, floatNxM_t vm); as defined in ARM-software/acle#324 --------- Co-authored-by: Hassnaa Hamdi <[email protected]>

This patch adds these intrinsics: // Variants are also available for: _s8 svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0"); according to PR#324[1] [1]ARM-software/acle#324

sallyarmneale

LGTM

…#97755) This patch was reverted because of a failing C test. It now has being solved and can be merged into main again This patch adds these intrinsics: // Variants are also available for: _s8 svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0"); according to PR#324[1] [1]ARM-software/acle#324 OBS.: Fix the clang test run line

…#97755) This patch was reverted because of a failing C test. It now has being solved and can be merged into main again This patch adds these intrinsics: // Variants are also available for: _s8 svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0"); according to PR#324[1] [1]ARM-software/acle#324 OBS.: Fix the clang test run line Address comments about the functions SelectMultiVectorLuti

This patch adds these intrinsics: // Variants are also available for: // [_s8], [_u16], [_s16], [_u32], [_s32], [_u64], [_s64] // [_bf16], [_f16], [_f32], [_f64] void svwrite_lane_zt[_u8](uint64_t zt0, svuint8_t zt, uint64_t idx) __arm_streaming __arm_inout("zt0"); void svwrite_zt[_u8](uint64_t zt0, svuint8_t zt) __arm_streaming __arm_inout("zt0"); according to PR#324[1] [1]ARM-software/acle#324

…) (#109953) This patch was reverted because of a failing C test. It now has being solved and can be merged into main again This patch adds these intrinsics: // Variants are also available for: _s8 svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0"); according to PR#324[1] [1]ARM-software/acle#324

…55) (#109953) This patch was reverted because of a failing C test. It now has being solved and can be merged into main again This patch adds these intrinsics: // Variants are also available for: _s8 svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0"); according to PR#324[1] [1]ARM-software/acle#324

This patch adds these intrinsics: // Variants are also available for: // [_s8], [_u16], [_s16], [_u32], [_s32], [_u64], [_s64] // [_bf16], [_f16], [_f32], [_f64] void svwrite_lane_zt[_u8](uint64_t zt0, svuint8_t zt, uint64_t idx) __arm_streaming __arm_inout("zt0"); void svwrite_zt[_u8](uint64_t zt0, svuint8_t zt) __arm_streaming __arm_inout("zt0"); according to PR#324[1] [1]ARM-software/acle#324

…55) (#109953) This patch was reverted because of a failing C test. It now has being solved and can be merged into main again This patch adds these intrinsics: // Variants are also available for: _s8 svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0"); according to PR#324[1] [1]ARM-software/acle#324

This patch adds these intrinsics: // Variants are also available for: // [_s8], [_u16], [_s16], [_u32], [_s32], [_u64], [_s64] // [_bf16], [_f16], [_f32], [_f64] void svwrite_lane_zt[_u8](uint64_t zt0, svuint8_t zt, uint64_t idx) __arm_streaming __arm_inout("zt0"); void svwrite_zt[_u8](uint64_t zt0, svuint8_t zt) __arm_streaming __arm_inout("zt0"); according to PR#324[1] [1]ARM-software/acle#324

…97602) This patch adds these intrinsics: // Variants are also available for: // [_s8], [_u16], [_s16], [_u32], [_s32], [_u64], [_s64] // [_bf16], [_f16], [_f32], [_f64] void svwrite_lane_zt[_u8](uint64_t zt0, svuint8_t zt, uint64_t idx) __arm_streaming __arm_inout("zt0"); void svwrite_zt[_u8](uint64_t zt0, svuint8_t zt) __arm_streaming __arm_inout("zt0"); according to PR#324[1] [1]ARM-software/acle#324

Lukacma reviewed Jun 21, 2024

View reviewed changes

neon_intrinsics/advsimd.md Outdated Show resolved Hide resolved

Lukacma mentioned this pull request Jun 27, 2024

[AArch64][NEON] Add intrinsics for LUTI llvm/llvm-project#96883

Merged

Lukacma reviewed Jun 27, 2024

View reviewed changes

main/acle.md Show resolved Hide resolved

rsandifo-arm reviewed Jun 28, 2024

View reviewed changes

main/acle.md Outdated Show resolved Hide resolved

Lukacma mentioned this pull request Jun 28, 2024

[AARCH64][SVE] Add intrinsics for SVE LUTI instructions llvm/llvm-project#97058

Merged

CarolineConcatto reviewed Jul 3, 2024

View reviewed changes

main/acle.md Outdated Show resolved Hide resolved

CarolineConcatto mentioned this pull request Jul 3, 2024

[Clang][LLVM][AArch64] Add intrinsic for MOVT SME2 instruction llvm/llvm-project#97602

Merged

momchil-velikov commented Jul 4, 2024

View reviewed changes

CarolineConcatto mentioned this pull request Jul 4, 2024

[Clang][LLVM][AArch64] Add intrinsic for LUTI4 SME2 instruction llvm/llvm-project#97755

Merged

andrewcarlotti reviewed Jul 8, 2024

View reviewed changes

tools/intrinsic_db/advsimd.csv Outdated Show resolved Hide resolved

andrewcarlotti reviewed Jul 8, 2024

View reviewed changes

tools/intrinsic_db/advsimd.csv Outdated Show resolved Hide resolved

tools/intrinsic_db/advsimd.csv Outdated Show resolved Hide resolved

This was referenced Jul 16, 2024

[AArch64] Implement NEON vamin/vamax intrinsics llvm/llvm-project#99041

Merged

[AArch64] Implement intrinsics for SVE FAMIN/FAMAX llvm/llvm-project#99042

Merged

momchil-velikov mentioned this pull request Jul 16, 2024

[AArch64] Implement intrinsics for SME2 FAMIN/FAMAX llvm/llvm-project#99063

Merged

ktkachov reviewed Jul 23, 2024

View reviewed changes

main/acle.md Outdated Show resolved Hide resolved

ktkachov reviewed Jul 23, 2024

View reviewed changes

main/acle.md Show resolved Hide resolved

rsandifo-arm reviewed Jul 31, 2024

View reviewed changes

momchil-velikov mentioned this pull request Jul 31, 2024

FP8 ACLE specification #323

Merged

8 tasks

momchil-velikov force-pushed the faminmax-and-luti branch from c4388fa to a6261ca Compare August 1, 2024 09:30

momchil-velikov added 3 commits August 28, 2024 10:27

Intrinsics for absolute minimum and maximum, and table lookup

374eb18

[fixup] Add lane/laneq to some intrinsics, use imm_idx consiste…

d9d0080

…ntly

[fixup] Replace svmovt_zt with svwrite_zt

8a17a84

sallyarmneale reviewed Sep 25, 2024

View reviewed changes

CarolineConcatto mentioned this pull request Sep 25, 2024

[Clang][LLVM][AArch64] Add intrinsic for LUTI4 SME2 instruction (#97755) llvm/llvm-project#109953

Merged

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Intrinsics for absolute minimum and maximum, and table lookup #324

Intrinsics for absolute minimum and maximum, and table lookup #324

Uh oh!

momchil-velikov commented Jun 13, 2024

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

momchil-velikov Jul 4, 2024

Uh oh!

Uh oh!

andrewcarlotti left a comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

rsandifo-arm commented Jul 31, 2024

Uh oh!

vhscampos commented Jul 31, 2024

Uh oh!

vhscampos commented Aug 5, 2024

Uh oh!

sallyarmneale left a comment

Uh oh!

Uh oh!

Intrinsics for absolute minimum and maximum, and table lookup #324

Intrinsics for absolute minimum and maximum, and table lookup #324

Uh oh!

Conversation

momchil-velikov commented Jun 13, 2024

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

momchil-velikov Jul 4, 2024

Choose a reason for hiding this comment

Uh oh!

Uh oh!

andrewcarlotti left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

rsandifo-arm commented Jul 31, 2024

Uh oh!

vhscampos commented Jul 31, 2024

Uh oh!

vhscampos commented Aug 5, 2024

Uh oh!

sallyarmneale left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!