Skip to content

perf: vectorize small u8 table take with NEON - #9571

Merged
joseph-isaacs merged 6 commits into
developfrom
ji/small-u8-table-take-neon
Oct 5, 2026
Merged

joseph-isaacs merged 6 commits into
developfrom
ji/small-u8-table-take-neon

Conversation

@joseph-isaacs

Copy link
Copy Markdown
Contributor

Rationale

Low-cardinality dictionaries with u8 codes and one-byte values can decode through an in-register NEON table instead of arbitrary scalar lookups.

Stacked on benchmark baseline #9566. The x86 implementation is intentionally split into the next PR.

Changes

  • Use NEON TBL on little-endian AArch64.
  • Apply the path to u8 codes, at most 16 one-byte values, and at least 64 rows.
  • Retain bounds checks and the existing fallback.
  • Add correctness and out-of-bounds tests.

CodSpeed wall-time results

Medians over 1,000 samples on Graviton3 metal, compared with #9566:

Rows Baseline NEON Speedup Time reduction
1M 552.4 µs 51.06 µs 10.82× 90.8%
16M 8.786 ms 799.7 µs 10.99× 90.9%

Runs: baseline, optimized.

Checks

  • focused small-table correctness and bounds tests
  • targeted Clippy with warnings denied
  • cargo +nightly fmt --all --check
  • git diff --check

Loading
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/performance A performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants