Skip to content

Canonicalize chunked nested types through the builder [builders-child-stack] - #8967

Open
robert3005 wants to merge 22 commits into
developfrom
claude/chunked-canonical-via-builder-9ze0t6
Open

robert3005 wants to merge 22 commits into
developfrom
claude/chunked-canonical-via-builder-9ze0t6

Conversation

@robert3005

@robert3005 robert3005 commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor

Always use builders when canonicalizing chunked arrays

@codspeed

codspeed Bot commented Jul 25, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 2 benchmarks

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 28 improved benchmarks
❌ 2 regressed benchmarks
✅ 2067 untouched benchmarks
🆕 10 new benchmarks
⏩ 503 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Simulation take[value_shape/duplicate99/struct8/nonnull/chunks=16/indices=100] 169.1 µs 200 µs -15.44%
❌ WallTime dict_canonicalize_gt_u8_neon[1000000] 487.9 µs 561.2 µs -13.06%
⚡ Simulation canonicalize[1024, 32] 52.5 µs 36.4 µs +44.24%
⚡ Simulation canonicalize[256, 32] 52.6 µs 36.5 µs +44.17%
⚡ Simulation chunked_opt_bool_into_canonical[(10, 100)] 277.7 µs 199 µs +39.56%
⚡ Simulation canonicalize[1024, 8] 38 µs 28.2 µs +34.85%
⚡ Simulation canonicalize[256, 8] 37.9 µs 28.2 µs +34.73%
⚡ Simulation canonicalize[16, 8] 38 µs 28.2 µs +34.66%
⚡ Simulation canonicalize[16, 32] 66.4 µs 50.5 µs +31.69%
⚡ Simulation canonicalize[256, 2] 35.3 µs 26.9 µs +31.02%
⚡ Simulation canonicalize[1024, 2] 35.2 µs 26.9 µs +30.82%
⚡ Simulation chunked_varbinview_into_canonical[(10, 100)] 367.6 µs 295.3 µs +24.52%
⚡ Simulation chunked_opt_bool_into_canonical[(100, 50)] 248.3 µs 202.2 µs +22.8%
⚡ Simulation canonicalize[16, 2] 44.6 µs 37.3 µs +19.77%
⚡ Simulation chunked_varbin_into_canonical[(10, 100)] 428.2 µs 358 µs +19.6%
⚡ Simulation chunked_varbinview_opt_into_canonical[(10, 100)] 553.8 µs 474.7 µs +16.66%
⚡ Simulation chunked_opt_bool_into_canonical[(1000, 10)] 101 µs 86.9 µs +16.22%
⚡ Simulation take[value_shape/shuffled/struct8/nonnull/chunks=16/indices=100] 3.3 ms 2.8 ms +15.38%
⚡ Simulation take_chunked_fsl_sorted[8, 64] 265 µs 233.2 µs +13.67%
⚡ Simulation take_chunked_fsl_sorted[32, 64] 279 µs 245.8 µs +13.54%
... ... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing claude/chunked-canonical-via-builder-9ze0t6 (e0929fb) with develop (900c1f2)

Open in CodSpeed

Footnotes

  1. 503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@claude claude Bot changed the title refactor(array): canonicalize chunked structs and FSLs through the builder refactor(array): canonicalize chunked structs and FSLs through the builder [builders-child-stack] Jul 25, 2026
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch 2 times, most recently from afed52f to 79f350a Compare July 29, 2026 07:35
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch 2 times, most recently from 1145029 to 8f39b06 Compare July 29, 2026 14:19
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from 8f39b06 to 825ab2e Compare July 31, 2026 13:35
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from 825ab2e to ce2eb1d Compare July 31, 2026 15:29
Comment thread vortex-array/src/arrays/chunked/vtable/canonical.rs Outdated
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from ce2eb1d to cefbc8e Compare July 31, 2026 16:54
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from cefbc8e to 0d4eb9a Compare July 31, 2026 17:22
@robert3005 robert3005 changed the title refactor(array): canonicalize chunked structs and FSLs through the builder [builders-child-stack] Canonicalize chunked nested types through the builder [builders-child-stack] Jul 31, 2026
@robert3005 robert3005 added the changelog/chore A trivial change label Jul 31, 2026
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from 0d4eb9a to 1e8346a Compare July 31, 2026 18:39
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from 1e8346a to 0431fa7 Compare July 31, 2026 18:43
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from 7fecff2 to 9a1bd56 Compare September 16, 2026 13:36
@robert3005

Copy link
Copy Markdown
Contributor Author

The goal here is to support going through execution loop for each of the chunks so it's not recursive execution. I am improving the performance but I think this would be a slight regression in synthetic benchmarks since instruction count will be slightly larger.

@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from 9a1bd56 to e47af6e Compare September 16, 2026 16:32
@robert3005

Copy link
Copy Markdown
Contributor Author

I realised I was not doing the full change, now we are going through the executor and things need optimisations

@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from 39a4e35 to 0df41f5 Compare September 17, 2026 00:54
@robert3005

robert3005 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

Codspeed will always be unhappy, simulated execution doesn't cover the fact that some code will only ever run once for n iterations. there's something here...

@robert3005

Copy link
Copy Markdown
Contributor Author

This turned into a lot more complicated change that I would have wanted but I think it's right now. There are improvements to the core execution loop that I opened pr for independently of this pr and we should review them first.

In the end nested types require special path since the operation for them is trivial, i.e. we know we are doing a transpose/swizzle operation without any value copying. Thus we keep it but we integrate it into execution loop and return ExecutionResult iteratively instead of running recursively to completion, we also skip chunks that are already canonical

claude and others added 22 commits October 2, 2026 10:34
`pack_struct_chunks`, `swizzle_fixed_size_list_chunks` and
`swizzle_list_chunks` existed because `append_to_builder` used to decode a
builder's children: without them, canonicalizing a `ChunkedArray` would
concatenate every chunk's children instead of reusing them. Their doc
comments describe what the builders now do on their own, so all three are
dead weight - the generic builder path produces the same swizzled array,
with each chunk's children kept as chunks of the combined child.

Only `Variant` still needs a hand-written pack, because there is no variant
builder.

This drops the one canonicalization path that allocated its buffers through
the session allocator: `swizzle_list_chunks` allocated its offsets and sizes
with `ctx.allocator()`, whereas `builder_with_capacity_in` still ignores the
allocator it is handed, so chunked primitives, structs and FSLs already went
around it. `list_canonicalize_uses_memory_session_allocator` guarded that
one path and goes with it; restoring the property means teaching the
builders to allocate through a `HostAllocatorRef`, not keeping a bespoke
list swizzle alive.

Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Robert Kruszewski <robert@spiraldb.com>

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
I, Robert Kruszewski <github@robertk.io>, hereby add my Signed-off-by to this commit: 5c8401d

Signed-off-by: Robert Kruszewski <github@robertk.io>
`ListBuilder` and `ListViewBuilder` wrote one offset - and for list views one
size - per row when appending empty or null lists, pushing the same value `n`
times into a `PrimitiveBuilder`. `append_n_values` writes the run in one go.

This is not a rare path: a sparse list array with a null or empty fill calls
`append_nulls(count)` with the gap between patches, which can be millions of
rows. `append_nulls(1_000_000)` on a `ListViewBuilder` goes from 22.2ms to
979us.

Also let a builder be told how many arrays are about to be appended.
`Chunked::append_to_builder` hands over one array per chunk, and a nested
builder keeps each as a chunk of its child, so the chunk lists grew one push
at a time. `reserve_chunks` lets them grow once. In isolation this is worth
36% of `ChildBuilder::append_array` at 100k chunks; end to end it measures as
noise, since canonicalization is dominated by the per-chunk work around it.

Signed-off-by: Claude <noreply@anthropic.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
The vtable hands `append_to_builder` an `ArrayView<'_, V>` - two borrows,
`Copy` - while the typed builder methods took `&Array<V>`. Bridging the two
meant `array.into_owned()`, which clones the `ArrayRef`, so every appended
array paid an atomic refcount pair for a builder that only reads from it.

The builders read through `TypedArrayRef`, which `ArrayView` implements, so
taking the view directly needs no other change. `BuffersWithOffsets::
from_array` and `VarBinViewBuilder::push_view` follow their caller.

`buffer_utilizations` and the two helpers it uses were inherent to
`VarBinViewArray` and so unreachable from a view. They read nothing an
`ArrayView` cannot supply, so they move to a blanket `VarBinViewCompactExt`
over `TypedArrayRef<VarBinView>`; the compaction entry points stay inherent.

Signed-off-by: Claude <noreply@anthropic.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
Canonicalizing a chunked array through the builder now costs roughly what
the bespoke swizzle it replaced did, instead of ~11% more instructions.
Three things were being paid once per chunk that only need paying once per
appended array:

- `ValidityBuilder` recorded a run per chunk for validities that are
  uniform, so a 32-chunk non-nullable array left 32 runs for
  `Validity::concat` to collapse and for `finish_with_nullability` to walk.
  Adjacent uniform runs of the same kind now extend rather than push.

- `append_to_builder` re-derived and re-compared the builder's dtype, and
  re-checked the builder's length, for every chunk. `Chunked` now appends
  its chunks through an `append_to_builder_unchecked` that skips both: the
  enclosing call already checked the chunked array's dtype, which every
  chunk shares, and its length, which the chunks' lengths sum to.

- `ChildBuilder::append_array` compared dtypes on every appended chunk.
  `ChildBuilder` is `pub(crate)`, and each of its callers is a nested
  builder passing down a child of an array whose dtype was checked against
  that builder on the way in, so the child's dtype follows from the
  parent's. The comparison becomes a debug assertion, which is what the
  invariant is.

Measured with callgrind on a 32-chunk x 1000-list x 1024-element
`FixedSizeList` canonicalization, against the pre-builder swizzle:

  swizzle                12,610,357 Ir
  builder, before        14,022,180 Ir   +11.2%
  builder, after         12,791,687 Ir    +1.4%

Signed-off-by: Claude <noreply@anthropic.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
`ArrayView` is not in scope in `varbinview::compact`, so rustdoc could not
resolve the bare link and the workspace doc build failed under
`-D warnings`. Link through `crate::array::ArrayView`, matching how the
rest of the crate writes out-of-scope doc links.

Signed-off-by: Robert Kruszewski <robert@spiraldb.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Chunked appends its chunks through the checked `append_to_builder` again,
paying the builder dtype comparison and the length post-condition per chunk
rather than once for the whole chunked array.

Signed-off-by: Robert Kruszewski <robert@spiraldb.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
Canonicalizing an N-chunk array through `execute_until` runs N+2 iterations,
and four things were being paid on every one of them that only need paying
once. Measured on `chunked_fsl_canonicalize`, median times drop 30-42%, with
the gain growing in the chunk count:

  chunks     before      after
       2   1.770 us   1.218 us   -31%
       8   3.214 us   1.996 us   -38%
      32   8.535 us   4.979 us   -42%

The timings are flat in list size, which is the tell: this is per-chunk
bookkeeping, not data movement.

- `AnyCanonical::matches` asked twelve encodings `array.is::<V>()` in turn,
  each a virtual `as_any` call plus a `TypeId` comparison. It now compares
  the array's encoding id against the twelve canonical ids, interned once.

- The loop then ran that check twice per iteration. `execute::<Canonical>`
  enters with `M = AnyCanonical`, so while the stack is empty the target
  predicate *is* the canonical stop condition. A `Matcher::IS_ANY_CANONICAL`
  associated constant lets the root case answer both with one scan, resolved
  at monomorphization rather than by comparing function pointers.

- The builder was never told how many arrays were coming. `reserve_chunks`
  only ran in `Chunked::append_to_builder`, which `execute_until` bypasses
  now that `Chunked::execute` drives the builder through `AppendChild`, so
  every nested builder's chunk list grew incrementally: a 128-chunk array of
  8-field structs took 40 reallocations, now none.

- `Chunked` built a real empty array of its own dtype to carry the final
  `Done`, which `finalize_done` discards in favour of the builder. For a
  nested dtype that is a recursive construction of empty children, so
  `ExecutionResult::done_into_builder` now names that case and carries a
  null placeholder instead. That, plus the reserves, takes a 128-chunk
  8-field struct from 99 allocations to 81.

Also defers the executor's `expected_dtype` clone to debug builds, where
`finalize_done` is the only thing that reads it back.

Signed-off-by: Robert Kruszewski <robert@spiraldb.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
…r trait

`Matcher::IS_ANY_CANONICAL` put an executor implementation detail on a public
trait, where every implementor saw it and any of them could set it wrongly -
a `true` on a matcher that is not `AnyCanonical` would make the executor stop
early, which is a correctness bug rather than a performance one. The executor
now asks the question itself, comparing `TypeId::of::<M>()` against
`AnyCanonical` once per `execute_until` call. It benchmarks the same as the
constant did, and the trait goes back to what it was.

Also closes a hole in the id-based `AnyCanonical::matches`: an encoding id
does not identify a `VTable` type on its own, because `ForeignArray`,
`ScalarFn` and the Python vtable each return a per-instance `self.id` and
nothing stops one being registered under a canonical encoding's id. The id
scan is now only a fast reject, with a hit still confirmed by the downcast in
`try_match`. Without that, `matches` could answer yes where `try_match`
answers `None`, and `Canonical::execute` turns that disagreement into a panic.

Signed-off-by: Robert Kruszewski <robert@spiraldb.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
…he list

`AnyCanonical::try_match` offered the array to each canonical encoding in
turn, so a `FixedSizeList` array paid nine `as_opt` calls - each a virtual
`as_any` plus a `TypeId` comparison - to reach its own arm, and an `Extension`
array all twelve. Counting them over a chunked canonicalize showed 18
downcasts per execution against a single canonical array. Comparing the
interned encoding id picks the one encoding worth asking, leaving one
downcast to confirm it.

That confirmation is what the previous commit paid for with a second walk of
the list. It is still needed, because an encoding id does not identify a
`VTable` type on its own, but it now costs one downcast rather than nine.

`matches` keeps its own expansion rather than delegating to `try_match`:
returning `bool` keeps the rejection path clear of the `CanonicalView` that
the `Option` would carry, and rejecting with a single `||` reduction over the
ids keeps it branchless. Both were worth about 4% on their own.

Measured on `chunked_fsl_canonicalize` at list size 1024, medians of three
runs, against the pre-optimization baseline re-measured in the same
conditions:

  chunks     before      after
       2   1.796 us   1.243 us   -31%
       8   3.219 us   2.069 us   -36%
      32   8.612 us   5.107 us   -41%

Both halves of the matcher are generated from one list of
`field => vtable => variant` triples, so they cannot disagree about which
encodings are canonical, and each arm's id comes from the same vtable that
arm downcasts to. The pairing of a vtable with its `CanonicalView` variant is
enforced by the compiler, since the view borrows `ArrayView<'_, V>`. A test
covers the remaining gap: that every canonical encoding appears in the list at
all, checked through `matches` and `try_match` alike.

Signed-off-by: Robert Kruszewski <robert@spiraldb.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
`data_mut` reached its own encoding data through `as_any_mut().downcast_mut()`,
spending a virtual call and a `TypeId` comparison to learn what `Array<V>`
already guarantees. The shared path beside it, `downcast_inner`, does not: it
asserts the invariant in debug and casts. This makes the two consistent.

The executor pays this once per chunk, through
`Chunked::execute` -> `with_next_builder_slot`, which threads the builder slot
counter through the array itself. Callgrind over a 32-chunk FSL canonicalize
attributed 1,977 instructions per execution to that function, a third of them
in `core::any`. It is now 1,417, and the whole canonicalize drops from 41,071
instructions to 40,314.

`as_any_mut` had no other callers, so `DynArrayData` loses it.

Wall clock on `chunked_fsl_canonicalize` at list size 1024, medians of three
runs, against the pre-optimization baseline measured in the same conditions:

  chunks     before      after
       2   1.796 us   1.227 us   -32%
       8   3.219 us   1.989 us   -38%
      32   8.612 us   4.999 us   -42%

Signed-off-by: Robert Kruszewski <robert@spiraldb.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
@robert3005
robert3005 force-pushed the claude/chunked-canonical-via-builder-9ze0t6 branch from d3e1f15 to e0929fb Compare October 2, 2026 09:35

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/chore A trivial change

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants