Skip to content

Compress fixed-size lists without their null rows - #10425

Open
robert3005 wants to merge 1 commit into
rk/compress-explicit-fallbackfrom
rk/compress-fsl-validity
Open

robert3005 wants to merge 1 commit into
rk/compress-explicit-fallbackfrom
rk/compress-fsl-validity

Conversation

@robert3005

@robert3005 robert3005 commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Add SparseScheme for FixedSizedList. Instead of compressing just the elements we attempt compressing whole lists using SparseArray

Readers need no changes, because Sparse already canonicalizes fixed-size lists.

On the data from #10417:

Before After
Mostly-null file (1 row in 20 valid) 25,296 B 9,056 B
Rewrite of both files 292,808 B 190,768 B

A fixed-size list stores `list_size` elements for every row, null or not, and
the compressor only ever compressed those elements, so it had to encode
whatever sits under the null rows. A UUID column that is mostly null was
compressed as dense bytes unless the elements happened to be 90% zeros, and
once mixed with a dense chunk not at all.

The compressor now offers fixed-size lists to schemes before compressing their
elements, after canonicalizing the elements so a scheme's output is compared
against the list's real size. The new `FixedSizeListSparseScheme` estimates the
saving of storing only the valid rows plus their bit-packed positions, and
when it is at least 10% encodes the list as `Sparse` with a null fill, its
values being the valid rows compressed recursively.

On the data from #10417, the mostly-null file shrinks from 25,296 to 9,056
bytes and the rewrite of both files from 292,808 to 190,768 bytes.

Signed-off-by: Claude <noreply@anthropic.com>

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XzXtnsrnEsBNVa9R2VqKQ
@codspeed

codspeed Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 1 benchmark

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 12 benchmarks spent significant time in system calls

System calls cannot be consistently instrumented, so they are not included in the measure, which understates the real cost. Please switch to the Walltime instrument to accurately measure system calls.

Measurement and system calls

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
❌ 1 regressed benchmark
✅ 2181 untouched benchmarks
⏩ 359 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ WallTime bitpack_blocked_compress_avx2 6.8 µs 7.6 µs -11.65%
⚡ ⚠️ Simulation decode_primitives[i64, (1000, 32)] 38.4 µs 24.8 µs +54.85%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing rk/compress-fsl-validity (c05ea8e) with develop (3d12093)2

Open in CodSpeed

Footnotes

  1. 359 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on rk/compress-explicit-fallback (c3ad942) during the generation of this report, so develop (3d12093) was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/feature A new feature

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants