Repository navigation
Conversation
Merging this PR will degrade performance by 24.89%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | decode_primitives[i64, (1000, 32)] |
24.1 µs | 38 µs | -36.41% |
| ❌ | WallTime | bitpack_blocked_compress_avx2 |
6.8 µs | 7.6 µs | -11.28% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing mk/bitpacked-v2-scheme (f638117) with mk/bitpacked-v2-wire (fc9ce15)
Footnotes
-
409 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
…ked.v2 is allowed Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
c98cb99 to
f638117
Compare
Tracking Issue: #10167
Stacked on #10244; this PR targets
mk/bitpacked-v2-wire.Summary
BitPackingSchemepacks each 1024-element block at its own bit width when the compressor's allowed serialized IDs includefastlanes.bitpacked.v2, followingFoRScheme(#10136). Default writes are unchanged: the core edition allows onlyfastlanes.bitpacked, sorefinekeeps v1.BitPackingSchemeBitPackingSchemegains v1 and v2 modes withv1()andv2()constructors, andDefaultis v1. The session registersBITPACKING_V1.refinereturns v2 whenfastlanes.bitpacked.v2is allowed, and v1 otherwise.produced_encodings: v1 declaresfastlanes.bitpacked; v2 declaresfastlanes.bitpacked.v2.FoRSchemev2 always encodes per-chunk references. It chooses every block's width withbitpack_to_best_bit_widths(BlockedBitPacked: encode blocked bitwidths #10203), even when every block chooses the same width, and compresses the block offsets as child 0 withcompress_child. Patches are handled as in v1: compressed in place, or moved intoPatchedunder experimental patches.FoRSchemestill bit-packs its encoded child withBITPACKING_V1, so FoR arrays keep a global width for now.CUDA
The CUDA preset excludes
fastlanes.bitpacked.v2, as it doesfastlanes.for.v2. CUDA decoding returns an error for per-block bit widths.Testing
refinepicks v2 exactly when the v2 ID is allowed, from either mode.tests/bitpacking_config.rs: with only BitPacking and every ID allowed, both blocks needing 1 to 8 bits and uniform 7-bit blocks serialize as v2, and nullable values round trip. The core edition never writes v2, and the CUDA preset writes v1.Follow-ups
FoRSchemewould bit-pack with v2 whenfastlanes.bitpacked.v2is allowed, so per-chunk references and per-block widths combine.FoRSchemeto per-chunk references whenfastlanes.for.v2is allowed #10136.