Skip to content

Don't pass extension arrays through the compressor unchanged - #10424

Open
robert3005 wants to merge 2 commits into
developfrom
rk/compress-explicit-fallback
Open

robert3005 wants to merge 2 commits into
developfrom
rk/compress-explicit-fallback

Conversation

@robert3005

@robert3005 robert3005 commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

The compressor indicates whether the array was returned unchanged. Extension scheme canonicalises input and only keeps extension scheme compressed arrays if they actually compressed anything instead of returning input unchanged

fix #10417

claude added 2 commits October 9, 2026 18:57
When no scheme compresses a canonical extension array, `choose_and_compress`
returns the input itself, and the extension branch kept it whenever its
`nbytes` beat the compressed storage. Canonicalizing an extension array only
canonicalizes the top level of its storage, so the input can still hold lazy
arrays such as `vortex.slice`, which cannot be serialized. A UUID column
mixing a dense chunk with a slice of a sparse one hit this on write
(#10417).

Use the compressed storage whenever no scheme produced a new array.

Signed-off-by: Claude <noreply@anthropic.com>

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XzXtnsrnEsBNVa9R2VqKQ
`choose_and_compress` returned its input whenever no scheme beat it, so every
caller got an `ArrayRef` that may or may not have been compressed. That is a
valid result for leaf arrays, but not for extension arrays, whose canonical
form leaves the storage below its top level lazy; the extension branch had to
detect the passthrough by pointer identity.

Return a `Selection` that says whether anything was compressed, so each caller
picks its fallback explicitly and the extension branch falls back to its
compressed storage by construction.

Signed-off-by: Claude <noreply@anthropic.com>

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XzXtnsrnEsBNVa9R2VqKQ
@codspeed

codspeed Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 1 benchmark

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 1 benchmark measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. This result is not comparable, so it counts as unchanged.

Preventing compiler optimizations

⚠️ 14 benchmarks spent significant time in system calls

System calls cannot be consistently instrumented, so they are not included in the measure, which understates the real cost. Please switch to the Walltime instrument to accurately measure system calls.

Measurement and system calls

⚡ 2 improved benchmarks
❌ 1 regressed benchmark
✅ 2180 untouched benchmarks
⏩ 359 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ WallTime bitpack_blocked_compress_avx2 6.7 µs 7.6 µs -11.45%
⚡ ⚠️ Simulation density_sweep_dense_runs[0.001] 47.8 µs 29.7 µs +60.79%
⚡ ⚠️ Simulation filter_powerlaw_by_mostly_true[250000] 150.2 µs 103.7 µs +44.77%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing rk/compress-explicit-fallback (c3ad942) with develop (731a231)

Open in CodSpeed

Footnotes

  1. 359 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/fix A bug fix

Projects

None yet

2 participants