Repository navigation
Optimise bitpacked filtering for valid runs and buffer filtering - #10281
robert3005 wants to merge 10 commits into
Conversation
…rializing indices BitPacked's filter kernel only handled very sparse masks (below 3-9% density), and for those it materialized the mask's `usize` indices first; any denser mask unpacked the whole array before filtering. The kernel now walks the selection one 1024-value FastLanes chunk at a time: from the mask's cached slices when it has them (as FixedSizeList element masks do) and otherwise straight from the bitmap. Empty chunks are skipped, fully selected chunks unpack directly into the output, chunks with few selected values use `unchecked_unpack_indices` on a small stack array, and the rest unpack into an L1-resident scratch chunk that is compacted with run copies, a branch-free loop or a trailing-zeros walk per mask word. Dense masks with scattered values still decline the kernel, so the vectorized canonical filter handles them. Adds a `bitpacking_filter` benchmark covering primitive i32 and FSL<i32> elements bit-packed to 16 bits. Each iteration builds a fresh mask so cached indices are not reused across iterations. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GEPn56FBLoa8Yq6jexAwWL
Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GEPn56FBLoa8Yq6jexAwWL
0e06ee5 to
fd431ca
Compare
Merging this PR will improve performance by 14.57%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ |
Simulation | random_i128[0.01] |
8 µs | 6.6 µs | +19.89% |
| ⚡ | Simulation | random_i128[0.8] |
175.9 µs | 154.8 µs | +13.65% |
| ⚡ | Simulation | random_i128[0.5] |
117.3 µs | 103.5 µs | +13.41% |
| ⚡ | Simulation | patterns_i128[Random] |
117.3 µs | 103.6 µs | +13.25% |
| ⚡ | Simulation | patterns_i128[Runs] |
118.1 µs | 104.7 µs | +12.8% |
| 🆕 | Simulation | fsl_i32[16, 0.1] |
N/A | 591.9 µs | N/A |
| 🆕 | Simulation | fsl_i32[4, 0.1] |
N/A | 813.3 µs | N/A |
| 🆕 | Simulation | primitive_i32[0.01] |
N/A | 179.4 µs | N/A |
| 🆕 | Simulation | primitive_i32[0.1] |
N/A | 519 µs | N/A |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing rk/bitpacked-filter (3b7d6ec) with develop (731a231)
Footnotes
-
359 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
…uses SIMD The chunked BitPacked filter declined only above a fixed 0.2 density, which only matches the canonical filter's SIMD compress crossover for 4-byte values. For u8 and u16 the canonical filter compacts with AVX-512/AVX2/NEON at much lower densities, so the chunked kernel was 2-6x slower there. Expose `uses_simd_compress::<T>(mask)` from vortex-array and decline whenever the canonical filter would compact with SIMD, except for very sparse masks (<= 1%) and masks with long cached runs. Also drop the branch-free word compaction, which ran the full word up to the last selected value and was 20-25% slower than the trailing-zeros walk at the densities where it was used. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AazYKY45Y5yuV5NCYwqbNf
|
Do you have any numbers here? |
Previous code assumed indices but these are almost never cached. Instead we handle the slices (which can be produced by lists) and bit buffers which are the default