## Summary
The random-access benchmark only ever read its data from local NVMe,
which hides the cost that matters most for point lookups: how many bytes
and round trips a format needs against an object store. This adds an S3
variant of the same benchmark so Vortex, Parquet, and Lance random
access can be tracked against remote storage, both on `develop` and on
demand for PRs.
The synthetic Parquet files keep the parquet-rs default row group and
page sizes; only the codec (zstd level 3, from #10157) is set.
## Changes
- `random-access-bench` gains `--remote-data-dir s3://bucket/prefix/`
and `--prepare-data`. Data is materialized locally as before, uploaded
verbatim, and then opened through `object_store` (Vortex and Parquet) or
Lance's own `s3://` provider. Remote measurements are named
`...-tokio-s3` and tagged with `s3` storage so they form a separate
series from the local-disk numbers. Arrow IPC has no object-store reader
and is skipped for remote runs.
- v3 ingest records have no storage field and their measurement ID
hashes only commit, dataset, format and open mode, so S3 runs use an
`-s3` dataset suffix (e.g. `taxi-s3/uniform`). This keeps them from
overwriting the local-disk rows without a schema change; local-disk IDs
are unchanged.
- `vortex-bench` adds `RemoteDataDir`, which maps a local data path to
its object key, plus `open_object_store` constructors for the Vortex and
Parquet accessors. The Parquet accessor keeps its cached footer and
re-opens through a boxed `AsyncFileReader` on each take, so the local
and object-store paths share one read path.
- CI: new `pr-bench-random-access-s3.yml` reusable workflow, wired into
`pr-bench-dispatch.yml` behind the `action/bench-random-access-s3` label
and `action/bench-all`. `pr-bench-runner.yml` learns `variant_id` and
`remote_data_dir` inputs, uploads the data before the run and deletes it
afterwards. `develop-bench.yml` gains a `Random Access (S3)` matrix
entry that refreshes
`s3://vortex-ci-benchmark-datasets/develop/random-access/` on every run.
- `lance-bench` enables Lance's `aws` feature so `Dataset::open` accepts
`s3://` URIs.
- Docs: README section for running against S3 and the new label in the
benchmarking guide.
Unit tests cover the key and URI mapping of `RemoteDataDir` and the
separate S3 ingest dataset.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01L627uApatcD62FGrQPutCL
---------
Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Co-authored-by: Claude <noreply@anthropic.com>
Summary
Split out of #9412 so the change to the benchmark data files lands on its own and the random-access baseline reset is attributed to it, rather than to the S3 variant.
The three synthetic random-access generators (
feature_vectors,nested_lists,nested_structs) passed no writer properties to parquet-rs, so their files were uncompressed. #10103 moved the other generators from Snappy to zstd level 3 but did not cover these, since they never set a codec at all.Changes
random_access_writer_propertiesinvortex-bench, which sets zstd level 3 and leaves row group and page sizes at the parquet-rs defaults.The Parquet-to-Vortex conversion is unaffected, so the derived Vortex files are byte-identical.