Skip to content

Release 0.3.0: canonical Arrow spine, cross-implementation conformance, 435 datasets - #20

Merged
mprammer merged 1 commit into
developfrom
mp/v0.3.0
Sep 28, 2026
Merged

mprammer merged 1 commit into
developfrom
mp/v0.3.0

Conversation

@mprammer

Copy link
Copy Markdown
Contributor

Cuts 0.3.0. Every build now writes a canonical Arrow IPC file and exports each format from it through a registered writer, and a conformance harness measures seven writer lanes against each other: pyarrow, arrow-rs, parquet-java through the parquet-arrow-java bridge (a submodule, pinned to its v0.2.0 release), Hardwood, and Vortex through Python, Rust and JNI. Its ledger ships in docs/v2/compliance.json; the only failures it records are two upstream Hardwood 1.1.0.Beta1 reader bugs. Every writer reads back what it writes, and a format a writer cannot produce is recorded with the measured error rather than a prose opt-out. The catalog grows from 250 to 435 datasets and gains versioned catalog bundles, native readers for Rust, C, C++ and Java, and the raincloud command.

This is a breaking release. The pipeline moves from scripts.pipeline to raincloud.pipeline, default data locations move to the platform's user directories, schema_version becomes 2 (outputs/v2, docs/v2), and reads no longer build by default; the Changed section of CHANGELOG.md is the upgrade guide. Tag v0.3.0 and the GitHub release follow on merge.

🤖 Generated with Claude Code

…e, 435 datasets

Bumps to 0.3.0. Every build now writes a canonical Arrow IPC file and
exports each format from it through a registered writer. Seven writer
lanes (pyarrow, arrow-rs, parquet-java through parquet-arrow-java,
Hardwood, and Vortex through Python, Rust and JNI) are measured against
each other by a conformance harness whose ledger ships in
docs/v2/compliance.json. Every writer reads back what it writes, and a
format a writer cannot produce is recorded with the measured failure
instead of a hand-written reason. The catalog grows from 250 to 435
datasets, including generated TPC-H and TPC-DS tables, and gains
versioned catalog bundles, native readers for Rust, C, C++ and Java,
and the raincloud command.

It is a breaking release: the pipeline moves from scripts.pipeline to
raincloud.pipeline, default data locations move to the platform's user
directories, schema_version becomes 2 (outputs/v2 and docs/v2), and
reads no longer build by default. Ingest no longer drops upstream
records or values: JSONBench rejoins the 16 records its dumps wrap, and
uci-thyroid-disease, uci-diabetes and uci-online-retail-ii keep what
they used to lose. CHANGELOG.md has the full list.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: Martin Prammer <martin@spiraldb.com>
@mprammer
mprammer marked this pull request as ready for review September 28, 2026 04:11
@mprammer
mprammer merged commit 63dd1d0 into develop Sep 28, 2026
12 checks passed
@mprammer
mprammer deleted the mp/v0.3.0 branch September 28, 2026 04:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant