Repository navigation
duckdb: support struct extract column indices #10297
Description
Activity
- addedchangelog/featureA new featureA new featuregood first issueGood for newcomersGood for newcomersext/duckdbRelates to the DuckDB integrationRelates to the DuckDB integration
on Oct 5, 2026 - added a parent issue
on Oct 5, 2026 I started on this and got steps 1–3 working on the scan side (paths across the FFI,
get_itemchain with the cast, no stats for extracted columns, plus the aggregation guard viaduckdb_reader_is_aggregate) — draft PR #10405.It does not land yet: with the callback enabled,
SELECT s.y, s.x.b, s.x.adies inside DuckDB withINTERNAL Error: ExpressionExecutor::Execute called with a result vector of type INTEGER that does not match expression type STRUCT(x STRUCT(a INTEGER, b INTEGER), y INTEGER) at MultiFileReader::FinalizeChunkMultiFileColumnMapper::MapColumnStructrebuilds the struct from the selected children and never checksIsPushdownExtract(), soFinalizeChunkapplies a struct expression to a vector typed as the extracted field.ColumnIndexType::PUSHDOWN_EXTRACTlooks like storage-table_scan-only in 1.5.5, and no MFR-based reader sets the callback.So the last piece is a DuckDB-side change in
multi_file_column_mapper.cpp. Is a patch undervortex-duckdb/patches/the route you want for that, or is there something an MFR can set to tell it the scan already returned the leaf? The draft has the scan side and a#[ignore]d e2e test that reproduces the failure, if it helps to judge.You are right, and I had the diagnosis before the conclusion the wrong way round. I went and looked at the pinned DuckDB, because that also answers where the retry has to happen.
On
develop's pin, DuckDB 1.5.5 (and the same holds for the newest release, v1.5.6):src/common/multi_file/multi_file_column_mapper.cpphas noIsPushdownExtract()at all, soMapColumnStructrebuilds the struct from the selected children andFinalizeChunkthen applies a struct expression to a vector sized for the leaf. That is exactly theINTERNAL ErrorI hit.extension/parquet/parquet_multi_file_info.cppnever setssupports_pushdown_extracteither, so on that pin there is no MFR-based reader to copy from.
At
myrrc/duckdb-2.0's pin, DuckDB561522aea, both are present:IsPushdownExtract()in the mapper, andsupports_pushdown_extract = ParquetScanSupportPushdownExtractin Parquet. So the temporary patch is what made this work on 1.5.5, and removing it in 2.0 is right because it is upstreamed. "The last piece is a patch underpatches/" was the 1.5.5 workaround, not the solution.Next on my side: rebase
feat/duckdb-struct-extract-pushdownontomyrrc/duckdb-2.0, re-enable the#[ignore]d nested test and see whether the scan side holds unchanged.Two questions, and I will follow whichever you prefer:
- Is
myrrc/duckdb-2.0(Duckdb 2.0 #9141) the base you meant, or should I wait for it to land ondevelopfirst? - If waiting, would you rather I keep [WIP] feat(duckdb): read pushed-down struct extracts #10405 as a draft on
developwith the callback off until the pin moves, or close it and open it fresh on the new base?
One thing that may be useful either way:
remove_unused_columns.cppdisables pushdown extract whenget.function.statisticsis set (statistics_extendedmust be used, orstatisticsleft NULL). We already usestatistics_extended, but it is a silent failure mode for any reader that sets both.Rebased onto
myrrc/duckdb-2.0and you were right. The nested test is un-ignored and green:cargo test -p vortex-duckdb --lib test result: ok. 216 passed; 0 failed; 3 ignored; 0 measured; 0 filtered outcargo fmt --checkis clean too.One part of the diff changed, and it got smaller rather than bigger: on 1.5.5 I had to override
fn.statistics_extendedto report no statistics for an extracted path. On 2.0 that override is redundant, because DuckDB's ownMultiFileScanStatsExtendedalready returns nothing for anIsPushdownExtract()column and for a column read as a type other than the stored one. So that piece is gone rather than reworked.One note for
#9141: it is 92 commits behinddevelop, so I took only my own commit withgit rebase --onto myrrc/duckdb-2.0 develop HEAD. A plaingit rebase myrrc/duckdb-2.0tries to replay all ofdevelop's commits onto the branch, which is the other direction.I have not pushed the rebased branch.
#10405can be updated as soon as#9141lands ondevelop, which is what I would default to. If you would rather review the diff against the 2.0 pin right now, I can retarget the PR atmyrrc/duckdb-2.0and force-push the rebased history instead. Which do you prefer?- Let's merge 2.0 changes. This will happen on Oct 21 or slightly later if Duckdb postpones the release. Then I'll be able to review and merge your change.
Duckdb has nested struct accessors like x.y.z. This generally produces a get_item expression which we handle in filters, so WHERE x.y.z = 1 will not read x into memory, but will read only x.y.z. However, projection pushdown doesn't support this, so SELECT x.y.z reads x.
Duckdb has another mechanism to optimize this which is a TableFunction callback supports_pushdown_extract. If you define it, you get special ColumnIndex with type STRUCT_EXTRACT.
In order to support this, we need to: