Repository navigation
perf: vectorize small u8 table take with AVX2 - #9572
Conversation
Merging this PR will regress 1 benchmark
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | take[small_m/shuffled/primitive/nonnull/chunks=2048/indices=64] |
280.1 µs | 481 µs | -41.76% |
| ⚡ | WallTime | dict_canonicalize_gt_u8_avx512[1000000] |
422.3 µs | 55.4 µs | ×7.6 |
| ⚡ | WallTime | dict_canonicalize_gt_u8_avx2[1000000] |
423.4 µs | 55.9 µs | ×7.6 |
| ⚡ | Simulation | decode_primitives[u8, (2000, 8)] |
44.5 µs | 33.5 µs | +32.83% |
| ⚡ | Simulation | decode_primitives[u8, (2000, 4)] |
44.5 µs | 33.5 µs | +32.83% |
| ⚡ | Simulation | decode_primitives[u8, (2000, 2)] |
44.5 µs | 33.5 µs | +32.83% |
| ⚡ | Simulation | decode_primitives[u8, (1000, 8)] |
37 µs | 31.5 µs | +17.45% |
| ⚡ | Simulation | decode_primitives[u8, (1000, 4)] |
37 µs | 31.5 µs | +17.45% |
| ⚡ | Simulation | decode_primitives[u8, (1000, 2)] |
37.4 µs | 31.8 µs | +17.43% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ji/small-u8-table-take-avx2 (25dce5c) with develop (e8e48c8)
Footnotes
-
518 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
|
This PR has been marked as stale because it has been open for 14 days with no activity. Please comment or remove the stale label if you wish to keep it active, otherwise it will be closed in 7 days |
21fa748 to
e97db31
Compare
0129184 to
134d1c5
Compare
|
This PR has been marked as stale because it has been open for 14 days with no activity. Please comment or remove the stale label if you wish to keep it active, otherwise it will be closed in 7 days |
Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
7d23aa5 to
25dce5c
Compare
Adds AVX2
VPSHUFBsmall-table lookup on top of the NEON implementation merged in #9571. It handlesu8indices, up to 16 one-byte values, and at least 64 rows, with runtime AVX2 detection and the existing gather/scalar fallbacks.The byte-value trait implementations live in the architecture-gated module. Table setup uses
copy_from_slice, scalar tails use safe writes, and output buffers preserve the caller's allocator. Bounds tests cover invalid indices in both vector and scalar-tail positions on all targets.For the 1M-row small-u8 dictionary benchmark, CodSpeed reports 55.4 µs versus 422.1 µs, a 7.6× speedup, on the AVX2 build. The AVX-512-enabled build runs the same AVX2 kernel and reports 55.2 µs. These measurements are from commit
d138ae05b5, before the subsequent merge ofdevelop.Benchmark comparison
Validation: all reported CI checks on
d138ae05b5passed apart from pending review policy. After mergingdevelop, the final diff was reviewed and new CI will validate the updated head; local tests were not rerun.