Repository navigation
perf: vectorize small u8 table take with NEON - #9571
Merged
Merged
Conversation
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale
Low-cardinality dictionaries with
u8codes and one-byte values can decode through an in-register NEON table instead of arbitrary scalar lookups.Stacked on benchmark baseline #9566. The x86 implementation is intentionally split into the next PR.
Changes
TBLon little-endian AArch64.u8codes, at most 16 one-byte values, and at least 64 rows.CodSpeed wall-time results
Medians over 1,000 samples on Graviton3 metal, compared with #9566:
Runs: baseline, optimized.
Checks
cargo +nightly fmt --all --checkgit diff --check