Repository navigation
perf: speed up runtime pruning and run-end execution - #10416
Draft
joseph-isaacs wants to merge 6 commits into
Draft
joseph-isaacs wants to merge 6 commits into
joseph-isaacs wants to merge 6 commits into
Conversation
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Apply supported join runtime filters before Vortex-to-Arrow conversion, and avoid decoding the same compressed run-end page at every binary-search comparison. Default Vortex retains the runtime-pruning gains while Compact recovers the regressions caused by repeated PCO decoding.
Changes
Sendfor the existing sorted-search adapter, with a compile-time regression test.Results
All 99 TPC-DS SF1 queries completed locally and on CI for the current
3850f40source versus exact base731a2314. Negative values mean lower geometric mean runtime. The CI report labels the overall result no clear signal, low confidence; these are raw hot-median comparisons.Parquet controls are within 0.3% on CI and 0.4% locally. Local measurements use 30 workers, 94 DataFusion partitions, ARM64, and exactly the same input files for both revisions; hot medians discard the first of eleven iterations. CI uses 94 CPUs, native CPU compilation, a cached baseline, and hot medians discarding the first of five iterations.
DataFusion Compact Q92 remains a regression: about 15% in the final local run, 13% in a separate 21-iteration repeat, and 25% in the current CI run. The full tables retain every query, including regressions. Paired Q92 profiles show additional PCO decoding and integer-membership work. Matching 250-iteration Q19 profiles isolate the retained-probe change: sampled CPU cycles fall about 59%, with approximately 81% less work in the largest PCO u32 decoding stage.
Validation
3850f40, including SQL logic tests, Rust linting, and integrations.Sendcontract, and run-end slicing.vortex-array,vortex-pco, andvortex-runend, all targets and features, with-D warnings; formatting usesnightly-2026-09-10.3850f40.All 99 CI query runtimes — default Vortex and Compact
TPC-DS SF1 — default Vortex and Compact, all 99 queries
PR
3850f40afa333abd3127f1358a4811131eddc3e3versus731a23142e9dc73345f7f79ecc8b665f1d49ba32. Official CI report.All runtimes are hot medians in milliseconds. Negative changes mean faster. Official verdict: No clear signal (low confidence).
Datafusion
Duckdb
All 13 SQL suites — original full run, hot runtime comparisons
Full SQL CI results for
3850f40afa333abd3127f1358a4811131eddc3e3versus731a23142e9dc73345f7f79ecc8b665f1d49ba32.Raw hot median geometric mean runtime changes; negative means faster. The official reports separately estimate confidence and environmental drift.
The FineWeb S3 DataFusion default outlier did not repeat. Both runs use the same source and cached base; both reports label the environment too noisy for confident attribution. The original full-run measurements remain above. The report comment now displays the targeted repeat.
These measurements show variability; they do not establish a confident S3 performance improvement.