Add Rv64v vlen128 NTT - #1875
Open
ishanrk wants to merge 3 commits into
Open
Conversation
Signed-off-by: Ishan Kumthekar <ishanrk@protonmail.com>
Signed-off-by: Ishan Kumthekar <ishanrk@protonmail.com>
Signed-off-by: Ishan Kumthekar <ishanrk@protonmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary:
add VLEN=128 support to the existing rv64v backend, including forward/inverse NTT and added the switch for VLEN=128 detection in meta.h where it is decided whether to run portable or RVV code.
The VLEN=128 implementation uses a separate NTT schedule (You can pretty much copy the VLEN=256 schedule until the last 3 layers of NTT when the distance between butterflies becomes 8, 4, and 2 and you use slightly different bitswap permutations with LMUL=4).
Do you expect this change to impact performance: Yes/No
YES ONLY FOR RISC-V VLEN = 128
If yes, please provide local benchmarking results.
On a Kendryte K230 the new NTT is roughly 1.7x faster than the existing portable C fallback (~41% fewer cycles).