Repository navigation
Canonicalize chunked nested types through the builder [builders-child-stack] - #8967
robert3005 wants to merge 22 commits into
Conversation
Merging this PR will regress 2 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | take[value_shape/duplicate99/struct8/nonnull/chunks=16/indices=100] |
169.1 µs | 200 µs | -15.44% |
| ❌ | WallTime | dict_canonicalize_gt_u8_neon[1000000] |
487.9 µs | 561.2 µs | -13.06% |
| ⚡ | Simulation | canonicalize[1024, 32] |
52.5 µs | 36.4 µs | +44.24% |
| ⚡ | Simulation | canonicalize[256, 32] |
52.6 µs | 36.5 µs | +44.17% |
| ⚡ | Simulation | chunked_opt_bool_into_canonical[(10, 100)] |
277.7 µs | 199 µs | +39.56% |
| ⚡ | Simulation | canonicalize[1024, 8] |
38 µs | 28.2 µs | +34.85% |
| ⚡ | Simulation | canonicalize[256, 8] |
37.9 µs | 28.2 µs | +34.73% |
| ⚡ | Simulation | canonicalize[16, 8] |
38 µs | 28.2 µs | +34.66% |
| ⚡ | Simulation | canonicalize[16, 32] |
66.4 µs | 50.5 µs | +31.69% |
| ⚡ | Simulation | canonicalize[256, 2] |
35.3 µs | 26.9 µs | +31.02% |
| ⚡ | Simulation | canonicalize[1024, 2] |
35.2 µs | 26.9 µs | +30.82% |
| ⚡ | Simulation | chunked_varbinview_into_canonical[(10, 100)] |
367.6 µs | 295.3 µs | +24.52% |
| ⚡ | Simulation | chunked_opt_bool_into_canonical[(100, 50)] |
248.3 µs | 202.2 µs | +22.8% |
| ⚡ | Simulation | canonicalize[16, 2] |
44.6 µs | 37.3 µs | +19.77% |
| ⚡ | Simulation | chunked_varbin_into_canonical[(10, 100)] |
428.2 µs | 358 µs | +19.6% |
| ⚡ | Simulation | chunked_varbinview_opt_into_canonical[(10, 100)] |
553.8 µs | 474.7 µs | +16.66% |
| ⚡ | Simulation | chunked_opt_bool_into_canonical[(1000, 10)] |
101 µs | 86.9 µs | +16.22% |
| ⚡ | Simulation | take[value_shape/shuffled/struct8/nonnull/chunks=16/indices=100] |
3.3 ms | 2.8 ms | +15.38% |
| ⚡ | Simulation | take_chunked_fsl_sorted[8, 64] |
265 µs | 233.2 µs | +13.67% |
| ⚡ | Simulation | take_chunked_fsl_sorted[32, 64] |
279 µs | 245.8 µs | +13.54% |
| ... | ... | ... | ... | ... | ... |
ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing claude/chunked-canonical-via-builder-9ze0t6 (e0929fb) with develop (900c1f2)
Footnotes
-
503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
afed52f to
79f350a
Compare
1145029 to
8f39b06
Compare
8f39b06 to
825ab2e
Compare
825ab2e to
ce2eb1d
Compare
ce2eb1d to
cefbc8e
Compare
cefbc8e to
0d4eb9a
Compare
0d4eb9a to
1e8346a
Compare
1e8346a to
0431fa7
Compare
7fecff2 to
9a1bd56
Compare
|
The goal here is to support going through execution loop for each of the chunks so it's not recursive execution. I am improving the performance but I think this would be a slight regression in synthetic benchmarks since instruction count will be slightly larger. |
9a1bd56 to
e47af6e
Compare
|
I realised I was not doing the full change, now we are going through the executor and things need optimisations |
39a4e35 to
0df41f5
Compare
|
|
|
This turned into a lot more complicated change that I would have wanted but I think it's right now. There are improvements to the core execution loop that I opened pr for independently of this pr and we should review them first. In the end nested types require special path since the operation for them is trivial, i.e. we know we are doing a transpose/swizzle operation without any value copying. Thus we keep it but we integrate it into execution loop and return ExecutionResult iteratively instead of running recursively to completion, we also skip chunks that are already canonical |
`pack_struct_chunks`, `swizzle_fixed_size_list_chunks` and `swizzle_list_chunks` existed because `append_to_builder` used to decode a builder's children: without them, canonicalizing a `ChunkedArray` would concatenate every chunk's children instead of reusing them. Their doc comments describe what the builders now do on their own, so all three are dead weight - the generic builder path produces the same swizzled array, with each chunk's children kept as chunks of the combined child. Only `Variant` still needs a hand-written pack, because there is no variant builder. This drops the one canonicalization path that allocated its buffers through the session allocator: `swizzle_list_chunks` allocated its offsets and sizes with `ctx.allocator()`, whereas `builder_with_capacity_in` still ignores the allocator it is handed, so chunked primitives, structs and FSLs already went around it. `list_canonicalize_uses_memory_session_allocator` guarded that one path and goes with it; restoring the property means teaching the builders to allocate through a `HostAllocatorRef`, not keeping a bespoke list swizzle alive. Signed-off-by: Claude <noreply@anthropic.com> Signed-off-by: Robert Kruszewski <robert@spiraldb.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
I, Robert Kruszewski <github@robertk.io>, hereby add my Signed-off-by to this commit: 5c8401d Signed-off-by: Robert Kruszewski <github@robertk.io>
`ListBuilder` and `ListViewBuilder` wrote one offset - and for list views one size - per row when appending empty or null lists, pushing the same value `n` times into a `PrimitiveBuilder`. `append_n_values` writes the run in one go. This is not a rare path: a sparse list array with a null or empty fill calls `append_nulls(count)` with the gap between patches, which can be millions of rows. `append_nulls(1_000_000)` on a `ListViewBuilder` goes from 22.2ms to 979us. Also let a builder be told how many arrays are about to be appended. `Chunked::append_to_builder` hands over one array per chunk, and a nested builder keeps each as a chunk of its child, so the chunk lists grew one push at a time. `reserve_chunks` lets them grow once. In isolation this is worth 36% of `ChildBuilder::append_array` at 100k chunks; end to end it measures as noise, since canonicalization is dominated by the per-chunk work around it. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
The vtable hands `append_to_builder` an `ArrayView<'_, V>` - two borrows, `Copy` - while the typed builder methods took `&Array<V>`. Bridging the two meant `array.into_owned()`, which clones the `ArrayRef`, so every appended array paid an atomic refcount pair for a builder that only reads from it. The builders read through `TypedArrayRef`, which `ArrayView` implements, so taking the view directly needs no other change. `BuffersWithOffsets:: from_array` and `VarBinViewBuilder::push_view` follow their caller. `buffer_utilizations` and the two helpers it uses were inherent to `VarBinViewArray` and so unreachable from a view. They read nothing an `ArrayView` cannot supply, so they move to a blanket `VarBinViewCompactExt` over `TypedArrayRef<VarBinView>`; the compaction entry points stay inherent. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
Canonicalizing a chunked array through the builder now costs roughly what the bespoke swizzle it replaced did, instead of ~11% more instructions. Three things were being paid once per chunk that only need paying once per appended array: - `ValidityBuilder` recorded a run per chunk for validities that are uniform, so a 32-chunk non-nullable array left 32 runs for `Validity::concat` to collapse and for `finish_with_nullability` to walk. Adjacent uniform runs of the same kind now extend rather than push. - `append_to_builder` re-derived and re-compared the builder's dtype, and re-checked the builder's length, for every chunk. `Chunked` now appends its chunks through an `append_to_builder_unchecked` that skips both: the enclosing call already checked the chunked array's dtype, which every chunk shares, and its length, which the chunks' lengths sum to. - `ChildBuilder::append_array` compared dtypes on every appended chunk. `ChildBuilder` is `pub(crate)`, and each of its callers is a nested builder passing down a child of an array whose dtype was checked against that builder on the way in, so the child's dtype follows from the parent's. The comparison becomes a debug assertion, which is what the invariant is. Measured with callgrind on a 32-chunk x 1000-list x 1024-element `FixedSizeList` canonicalization, against the pre-builder swizzle: swizzle 12,610,357 Ir builder, before 14,022,180 Ir +11.2% builder, after 12,791,687 Ir +1.4% Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
`ArrayView` is not in scope in `varbinview::compact`, so rustdoc could not resolve the bare link and the workspace doc build failed under `-D warnings`. Link through `crate::array::ArrayView`, matching how the rest of the crate writes out-of-scope doc links. Signed-off-by: Robert Kruszewski <robert@spiraldb.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
Chunked appends its chunks through the checked `append_to_builder` again, paying the builder dtype comparison and the length post-condition per chunk rather than once for the whole chunked array. Signed-off-by: Robert Kruszewski <robert@spiraldb.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
Canonicalizing an N-chunk array through `execute_until` runs N+2 iterations,
and four things were being paid on every one of them that only need paying
once. Measured on `chunked_fsl_canonicalize`, median times drop 30-42%, with
the gain growing in the chunk count:
chunks before after
2 1.770 us 1.218 us -31%
8 3.214 us 1.996 us -38%
32 8.535 us 4.979 us -42%
The timings are flat in list size, which is the tell: this is per-chunk
bookkeeping, not data movement.
- `AnyCanonical::matches` asked twelve encodings `array.is::<V>()` in turn,
each a virtual `as_any` call plus a `TypeId` comparison. It now compares
the array's encoding id against the twelve canonical ids, interned once.
- The loop then ran that check twice per iteration. `execute::<Canonical>`
enters with `M = AnyCanonical`, so while the stack is empty the target
predicate *is* the canonical stop condition. A `Matcher::IS_ANY_CANONICAL`
associated constant lets the root case answer both with one scan, resolved
at monomorphization rather than by comparing function pointers.
- The builder was never told how many arrays were coming. `reserve_chunks`
only ran in `Chunked::append_to_builder`, which `execute_until` bypasses
now that `Chunked::execute` drives the builder through `AppendChild`, so
every nested builder's chunk list grew incrementally: a 128-chunk array of
8-field structs took 40 reallocations, now none.
- `Chunked` built a real empty array of its own dtype to carry the final
`Done`, which `finalize_done` discards in favour of the builder. For a
nested dtype that is a recursive construction of empty children, so
`ExecutionResult::done_into_builder` now names that case and carries a
null placeholder instead. That, plus the reserves, takes a 128-chunk
8-field struct from 99 allocations to 81.
Also defers the executor's `expected_dtype` clone to debug builds, where
`finalize_done` is the only thing that reads it back.
Signed-off-by: Robert Kruszewski <robert@spiraldb.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
…r trait `Matcher::IS_ANY_CANONICAL` put an executor implementation detail on a public trait, where every implementor saw it and any of them could set it wrongly - a `true` on a matcher that is not `AnyCanonical` would make the executor stop early, which is a correctness bug rather than a performance one. The executor now asks the question itself, comparing `TypeId::of::<M>()` against `AnyCanonical` once per `execute_until` call. It benchmarks the same as the constant did, and the trait goes back to what it was. Also closes a hole in the id-based `AnyCanonical::matches`: an encoding id does not identify a `VTable` type on its own, because `ForeignArray`, `ScalarFn` and the Python vtable each return a per-instance `self.id` and nothing stops one being registered under a canonical encoding's id. The id scan is now only a fast reject, with a hit still confirmed by the downcast in `try_match`. Without that, `matches` could answer yes where `try_match` answers `None`, and `Canonical::execute` turns that disagreement into a panic. Signed-off-by: Robert Kruszewski <robert@spiraldb.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
…he list
`AnyCanonical::try_match` offered the array to each canonical encoding in
turn, so a `FixedSizeList` array paid nine `as_opt` calls - each a virtual
`as_any` plus a `TypeId` comparison - to reach its own arm, and an `Extension`
array all twelve. Counting them over a chunked canonicalize showed 18
downcasts per execution against a single canonical array. Comparing the
interned encoding id picks the one encoding worth asking, leaving one
downcast to confirm it.
That confirmation is what the previous commit paid for with a second walk of
the list. It is still needed, because an encoding id does not identify a
`VTable` type on its own, but it now costs one downcast rather than nine.
`matches` keeps its own expansion rather than delegating to `try_match`:
returning `bool` keeps the rejection path clear of the `CanonicalView` that
the `Option` would carry, and rejecting with a single `||` reduction over the
ids keeps it branchless. Both were worth about 4% on their own.
Measured on `chunked_fsl_canonicalize` at list size 1024, medians of three
runs, against the pre-optimization baseline re-measured in the same
conditions:
chunks before after
2 1.796 us 1.243 us -31%
8 3.219 us 2.069 us -36%
32 8.612 us 5.107 us -41%
Both halves of the matcher are generated from one list of
`field => vtable => variant` triples, so they cannot disagree about which
encodings are canonical, and each arm's id comes from the same vtable that
arm downcasts to. The pairing of a vtable with its `CanonicalView` variant is
enforced by the compiler, since the view borrows `ArrayView<'_, V>`. A test
covers the remaining gap: that every canonical encoding appears in the list at
all, checked through `matches` and `try_match` alike.
Signed-off-by: Robert Kruszewski <robert@spiraldb.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
`data_mut` reached its own encoding data through `as_any_mut().downcast_mut()`,
spending a virtual call and a `TypeId` comparison to learn what `Array<V>`
already guarantees. The shared path beside it, `downcast_inner`, does not: it
asserts the invariant in debug and casts. This makes the two consistent.
The executor pays this once per chunk, through
`Chunked::execute` -> `with_next_builder_slot`, which threads the builder slot
counter through the array itself. Callgrind over a 32-chunk FSL canonicalize
attributed 1,977 instructions per execution to that function, a third of them
in `core::any`. It is now 1,417, and the whole canonicalize drops from 41,071
instructions to 40,314.
`as_any_mut` had no other callers, so `DynArrayData` loses it.
Wall clock on `chunked_fsl_canonicalize` at list size 1024, medians of three
runs, against the pre-optimization baseline measured in the same conditions:
chunks before after
2 1.796 us 1.227 us -32%
8 3.219 us 1.989 us -38%
32 8.612 us 4.999 us -42%
Signed-off-by: Robert Kruszewski <robert@spiraldb.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZTu9XUrdRZZ91rx852DGf
Signed-off-by: Robert Kruszewski <github@robertk.io>
d3e1f15 to
e0929fb
Compare
Always use builders when canonicalizing chunked arrays