Skip to content

feat(layout): compile segment scans, filters, concatenations, and packs to pipelines - #10392

Open
joseph-isaacs wants to merge 8 commits into
ji/exec-2-exec-nodefrom
ji/exec-3-scan-concat-pack-nodes
Open

joseph-isaacs wants to merge 8 commits into
ji/exec-2-exec-nodefrom
ji/exec-3-scan-concat-pack-nodes

Conversation

@joseph-isaacs

@joseph-isaacs joseph-isaacs commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Stacked on #10391. Implements the pipeline compile hooks for the plans a struct of chunked columns lowers to, so such a plan runs end to end.

Changes

  • SegmentScan reads and decodes one segment and keeps the rows its split selects. A Filter over a scan compiles to the same source, and over any other child keeps the selected rows of each batch in a stage. A flat layout lowers to a bare segment scan, as on develop.
  • Concat builds only the chunks a split's mask selects a row of, so a chunk with no selected row is never read. It joins them with a source that drains its inlets in order. Rows inside one chunk are that chunk's rows, and no concatenation is built for them.
  • Pack zips its fields' streams into structs, slicing at the boundaries the fields share. A single field is wrapped as it passes.
  • Shared decoding. A segment several readers read is decoded once. One pass over the plan records the rows each segment is read for. The first reader builds one decoding pipeline that fans out into a port per reader, tagged by split, and a split's unclaimed ports are dropped when it finishes.

Tests port the node executor's cases with FIFO and LIFO delivery: views cutting every column at different places, unselected chunks never read, Pack streaming in row order, and a scan keeping its selection with and without a filter. They also check that several splits in one scan read each segment once, and that a finished scan holds no port.

Checks run: cargo clippy -p vortex-layout --all-targets --all-features -- -D warnings, cargo test -p vortex-layout --lib, cargo +nightly-2026-09-10 fmt -p vortex-layout.

🤖 Generated with Claude Code

https://claude.ai/code/session_01NWYmfHTFtaEdhdyd3u5hv3

@codspeed

codspeed Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 1 benchmark

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 1 benchmark measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. This result is not comparable, so it counts as unchanged.

Preventing compiler optimizations

⚠️ 13 benchmarks spent significant time in system calls

System calls cannot be consistently instrumented, so they are not included in the measure, which understates the real cost. Please switch to the Walltime instrument to accurately measure system calls.

Measurement and system calls

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
❌ 1 regressed benchmark
✅ 2181 untouched benchmarks
⏩ 359 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ WallTime bitpack_blocked_compress_avx2 6.7 µs 7.6 µs -11.98%
⚡ Simulation cold[(16, 64)] 356.2 µs 309.7 µs +15.02%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ji/exec-3-scan-concat-pack-nodes (a6f9c85) with ji/exec-2-exec-node (82fd9e5)2

Open in CodSpeed

Footnotes

  1. 359 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on ji/exec-2-exec-node (06f2233) during the generation of this report, so 3867f7b was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@joseph-isaacs
joseph-isaacs added this pull request to stack #10395 October 8, 2026 16:01
joseph-isaacs and others added 2 commits October 8, 2026 20:14
Add `Filter`, a plan operator with no predicate that keeps only the rows
of the selection it is executed with. Its child produces every row of the
same row domain, and values outside the selection are unspecified.

Flat layouts now lower to a `Filter` over their `SegmentScan`, so a bare
scan can stay dense and return every row of its range while the filter
above it keeps the selected ones. Plan display snapshots are updated for
the extra level.

Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NWYmfHTFtaEdhdyd3u5hv3
Implement `PlanVTable::exec` for the operators a projection of a struct
of chunked columns needs:

- `SegmentScanNode` requests its segment at start, decodes the delivered
  bytes, and emits the node's rows as one dense array. A `Filter` over a
  scan fuses into the same node with a filter mask, so the array holds
  exactly the selected rows. Decoded segments are shared through the
  graph's `DecodeCache`.
- `FilterNode` keeps the selected rows of any other child's dense arrays
  as they stream through, tracking its position with a cursor.
- `ConcatNode` spawns only the chunks overlapping its rows, one port
  each in chunk order, and emits the current chunk's rows as they
  arrive. Chunks whose reads land early wait in their ports.
- `PackNode` zips its fields as their rows arrive: whenever every port
  has rows, it takes as many as the shortest holds from each and emits
  one struct. Fields chunked alike stream one struct per chunk.

The layout tests drive lowered plans over in-memory segments with
scripted IO ordering and check what reaches the root. The scheduling
tests add hand-built sources under the real operators, including random
trees over random views with reads completing in random order, checked
against a model of the same tree.

Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NWYmfHTFtaEdhdyd3u5hv3
claude added 2 commits October 9, 2026 18:59
…ks to pipelines

Merges the pipeline runtime from ji/exec-2-exec-node and replaces this
branch's exec nodes with compile rules on the plans:

- SegmentScan reads and decodes one segment; a Filter over it keeps the
  selected rows in the same source, and over any other child keeps them
  in a stage per batch.
- Concat builds only the chunks a split's mask selects a row of, and
  joins them with a source that drains its inlets in order; rows inside
  one chunk are that chunk's rows.
- Pack zips its fields' streams into structs, slicing at the boundaries
  the fields share; a single field is wrapped as it passes.
- A segment more than one reader reads is decoded once: one pass over
  the plan records the rows each segment is read for, the first reader
  builds one decoding pipeline that fans out into a port per reader,
  tagged by split, and a split's unclaimed ports go when it finishes.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NWYmfHTFtaEdhdyd3u5hv3
…ep its own rows

A flat layout lowers to a bare segment scan again, as on develop. Every
plan now produces only the rows its split selects, a segment scan
included, so the scan's source applies the mask; a Filter over a scan
compiles to the same source.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NWYmfHTFtaEdhdyd3u5hv3
@joseph-isaacs joseph-isaacs changed the title feat(layout): add SegmentScan, Filter, Concat, and Pack exec nodes feat(layout): compile segment scans, filters, concatenations, and packs to pipelines Oct 9, 2026
claude added 4 commits October 9, 2026 19:46
Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NWYmfHTFtaEdhdyd3u5hv3
…akes

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NWYmfHTFtaEdhdyd3u5hv3
…y boundary

Pack used to emit a struct as soon as every field had a front batch, cut to the
shortest one. Fields chunked differently were sliced at every boundary of
every other field: 10,000 misaligned columns made 8,192 structs and
80 million slices per scan.

Now, when every field's front batch has the same length, Pack takes those
batches whole, so aligned fields still stream one struct per chunk with no
slicing. Otherwise it waits on the one field with the fewest queued rows until
that field is full or closed, then takes that many rows from every field. Whole
batches are joined as a ChunkedArray, not copied, and at most one batch per
field is sliced. Memory stays bounded by the inlet capacities.

Waiting costs O(1) per wake. Queued rows only grow until a struct is built, so
Pack resumes at the first inlet it found empty instead of rescanning, and while
filling it checks only the inlet it waits on.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NWYmfHTFtaEdhdyd3u5hv3
Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NWYmfHTFtaEdhdyd3u5hv3

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants