Skip to content

perf: speed up runtime pruning and run-end execution - #10416

Draft
joseph-isaacs wants to merge 6 commits into
developfrom
ji/tpcds-runtime-pruning-20261009
Draft

joseph-isaacs wants to merge 6 commits into
developfrom
ji/tpcds-runtime-pruning-20261009

perf: opt in to cached sorted probes for run ends

3850f40
Select commit
Loading
Failed to load commit list.
CodSpeed / CodSpeed Performance Analysis failed Oct 10, 2026

2 benchmarks regressed

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 15 benchmarks spent significant time in system calls

System calls cannot be consistently instrumented, so they are not included in the measure, which understates the real cost. Please switch to the Walltime instrument to accurately measure system calls.

Measurement and system calls

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 11 improved benchmarks
❌ 2 regressed benchmarks
✅ 2164 untouched benchmarks
🆕 167 new benchmarks
⏩ 365 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ WallTime bitpack_blocked_compress_avx2 6.7 µs 7.6 µs -11.34%
❌ Simulation and_false_constant 18.5 µs 20.6 µs -10.21%
⚡ Simulation i64_random_chunked[256] 15,817.7 µs 738.8 µs ×21
⚡ Simulation i64_random[256] 4,360.6 µs 247.3 µs ×18
⚡ Simulation decode_bool[10000_2_all_true] 175.7 µs 15.3 µs ×11
⚡ Simulation decode_bool[10000_2_all_false] 176.1 µs 15.8 µs ×11
⚡ Simulation decode_bool[10000_10_all_true] 47.5 µs 14.7 µs ×3.2
⚡ Simulation decode_bool[10000_10_all_false] 47.9 µs 15.2 µs ×3.1
⚡ Simulation copy_nullable[16384] 424.4 µs 313.6 µs +35.3%
⚡ Simulation decode_bool[10000_100_all_true] 17.9 µs 13.8 µs +29.45%
⚡ Simulation decode_bool[10000_100_all_false] 18.4 µs 14.4 µs +27.48%
⚡ Simulation sum_i64 225.1 µs 194.1 µs +15.98%
⚡ Simulation sum_v2_i64 224.8 µs 194.2 µs +15.76%
🆕 Simulation decode_primitive_array[i64, 1] N/A 2.9 ms N/A
🆕 Simulation decode_primitive_array[i64, 128] N/A 530.3 µs N/A
🆕 Simulation decode_primitive_array[i64, 16] N/A 693.5 µs N/A
🆕 Simulation decode_primitive_array[i64, 2] N/A 1.8 ms N/A
🆕 Simulation decode_primitive_array[i64, 4096] N/A 509.8 µs N/A
🆕 Simulation decode_primitive_array[u16, 1] N/A 2.3 ms N/A
🆕 Simulation decode_primitive_array[u16, 128] N/A 150.8 µs N/A
... ... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ji/tpcds-runtime-pruning-20261009 (3850f40) with develop (731a231)

Open in CodSpeed

Footnotes

  1. 365 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩