Skip to content

Optimise bitpacked filtering for valid runs and buffer filtering - #10281

Draft
robert3005 wants to merge 10 commits into
developfrom
rk/bitpacked-filter
Draft

robert3005 wants to merge 10 commits into
developfrom
rk/bitpacked-filter

Conversation

@robert3005

Copy link
Copy Markdown
Contributor

Previous code assumed indices but these are almost never cached. Instead we handle the slices (which can be produced by lists) and bit buffers which are the default

@robert3005 robert3005 added the changelog/performance A performance improvement label Oct 4, 2026
claude added 2 commits October 7, 2026 16:41
…rializing indices

BitPacked's filter kernel only handled very sparse masks (below 3-9% density),
and for those it materialized the mask's `usize` indices first; any denser mask
unpacked the whole array before filtering.

The kernel now walks the selection one 1024-value FastLanes chunk at a time:
from the mask's cached slices when it has them (as FixedSizeList element masks
do) and otherwise straight from the bitmap. Empty chunks are skipped, fully
selected chunks unpack directly into the output, chunks with few selected values
use `unchecked_unpack_indices` on a small stack array, and the rest unpack into
an L1-resident scratch chunk that is compacted with run copies, a branch-free
loop or a trailing-zeros walk per mask word. Dense masks with scattered values
still decline the kernel, so the vectorized canonical filter handles them.

Adds a `bitpacking_filter` benchmark covering primitive i32 and FSL<i32>
elements bit-packed to 16 bits. Each iteration builds a fresh mask so cached
indices are not reused across iterations.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GEPn56FBLoa8Yq6jexAwWL
Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GEPn56FBLoa8Yq6jexAwWL
@robert3005
robert3005 force-pushed the rk/bitpacked-filter branch from 0e06ee5 to fd431ca Compare October 7, 2026 15:41
@codspeed

codspeed Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Merging this PR will improve performance by 14.57%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 13 benchmarks spent significant time in system calls

System calls cannot be consistently instrumented, so they are not included in the measure, which understates the real cost. Please switch to the Walltime instrument to accurately measure system calls.

Measurement and system calls

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 5 improved benchmarks
✅ 2178 untouched benchmarks
🆕 4 new benchmarks
⏩ 359 skipped benchmarks1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
⚡ ⚠️ Simulation random_i128[0.01] 8 µs 6.6 µs +19.89%
⚡ Simulation random_i128[0.8] 175.9 µs 154.8 µs +13.65%
⚡ Simulation random_i128[0.5] 117.3 µs 103.5 µs +13.41%
⚡ Simulation patterns_i128[Random] 117.3 µs 103.6 µs +13.25%
⚡ Simulation patterns_i128[Runs] 118.1 µs 104.7 µs +12.8%
🆕 Simulation fsl_i32[16, 0.1] N/A 591.9 µs N/A
🆕 Simulation fsl_i32[4, 0.1] N/A 813.3 µs N/A
🆕 Simulation primitive_i32[0.01] N/A 179.4 µs N/A
🆕 Simulation primitive_i32[0.1] N/A 519 µs N/A

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing rk/bitpacked-filter (3b7d6ec) with develop (731a231)

Open in CodSpeed

Footnotes

  1. 359 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

…uses SIMD

The chunked BitPacked filter declined only above a fixed 0.2 density, which
only matches the canonical filter's SIMD compress crossover for 4-byte values.
For u8 and u16 the canonical filter compacts with AVX-512/AVX2/NEON at much
lower densities, so the chunked kernel was 2-6x slower there.

Expose `uses_simd_compress::<T>(mask)` from vortex-array and decline whenever
the canonical filter would compact with SIMD, except for very sparse masks
(<= 1%) and masks with long cached runs.

Also drop the branch-free word compaction, which ran the full word up to the
last selected value and was 20-25% slower than the trailing-zeros walk at the
densities where it was used.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AazYKY45Y5yuV5NCYwqbNf
@joseph-isaacs

Copy link
Copy Markdown
Contributor

Do you have any numbers here?

@joseph-isaacs
joseph-isaacs marked this pull request as draft October 8, 2026 12:57
Signed-off-by: Robert Kruszewski <github@robertk.io>
Comment thread encodings/fastlanes/src/bitpacking/compute/filter.rs Outdated
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Comment thread vortex-array/src/arrays/filter/execute/slice.rs Outdated
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/performance A performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants