Repository navigation
feat(bench): run the random access benchmark against S3 - #9412
Conversation
Merging this PR will not alter performance
|
Polar Signals Profiling ResultsLatest Run
Previous Runs (5)
Powered by Polar Signals Cloud |
Benchmarks: Random Access (S3) 📖Commits: PR No baseline is available for this benchmark yet; PR measurements are shown without comparison. random-access / vortex-file-compressed / ns (no group data, 0↑ 0↓)
random-access / parquet / ns (no group data, 0↑ 0↓)
random-access / lance / ns (no group data, 0↑ 0↓)
|
0fdd447 to
552c49a
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
This PR has been marked as stale because it has been open for 14 days with no activity. Please comment or remove the stale label if you wish to keep it active, otherwise it will be closed in 7 days |
8692404 to
42f46ab
Compare
|
This PR has been marked as stale because it has been open for 14 days with no activity. Please comment or remove the stale label if you wish to keep it active, otherwise it will be closed in 7 days |
Adds an S3 variant of the random-access benchmark. It is the same benchmark and the same code paths, only the data is read from an object store instead of local NVMe: - `--remote-data-dir s3://bucket/prefix/` opens the Vortex, Parquet and Lance files from S3, mirroring the local data directory layout; - `--prepare-data` materializes the local files (and prints their paths) so CI can upload them before the run; - remote measurements are suffixed `-tokio-s3` and reported with `s3` storage so they form a separate series from the local-disk numbers. CI runs it from a new `pr-bench-random-access-s3.yml` workflow (label `action/bench-random-access-s3`, also covered by `action/bench-all`) and from a new `Random Access (S3)` entry in the develop benchmark matrix. The shared PR benchmark runner gained `remote_data_dir` and `variant_id` inputs; it uploads the data to a per-run S3 prefix, runs the benchmark against it, and deletes the prefix afterwards. Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
`lance` is pinned with `default-features = false`, which drops the `aws`
feature it enables by default. Without it Lance has no `s3://` object store
provider registered and opening a remote dataset fails with:
Invalid user input: No object store provider found for scheme: 's3'
Parquet and Vortex already read from S3 fine; only Lance was affected.
Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
… files Every synthetic random-access dataset holds 1,000,000 rows and was written with default writer properties, whose `max_row_group_size` is 1Mi rows. A million rows never crosses that threshold, so each file is a single row group. Readers select row groups before rows, so a point lookup fetches and decodes the entire file: masked by page cache locally, ruinous over an object store, where one `take` on feature-vectors spends ~87s moving ~4GB. The tell is in the measurements: nested-lists and nested-structs take the same time under both access patterns to within 0.04%, because the indices never change what is read. Size row groups from the row width instead, targeting 128MiB and clamping to 8-64 batches of 1024 rows. The clamp is what fixes narrow rows: a pure byte budget would still leave nested-lists in one group. Sizes stay whole multiples of the 1024-row Arrow batches that `parquet_to_vortex_chunks` streams, so row group boundaries never split a batch and the derived Vortex files are unchanged -- only Parquet layout moves. Data pages also drop from 20k to 1024 rows, giving the page index resolution worth having here. Scan-oriented generators (TPC-H, SpatialBench, PolarSignals, ...) keep large row groups, which is right for full scans. This shifts Parquet random-access baselines once, most visibly on the correlated pattern. Uniform lookups still touch most row groups; only page-level row selection in the reader addresses those. Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
`set_max_row_group_size` is deprecated in favour of `set_max_row_group_row_count`, which takes an `Option<usize>` where `None` means unlimited. The deprecation warning fails the lint job. Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Rebasing onto develop picked up parquet 59, which deprecates `ParquetObjectReader` in favour of a hand-rolled `AsyncFileReader`. Keep the existing reader so the measured I/O path stays fixed against the numbers already collected, but scope the deprecation to its three uses rather than silencing warnings for the crate. Also moves `take_row_groups` above the test module that develop added, which `clippy::items_after_test_module` rejects. Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xmiM7RPQBBN9ycUXegLDv
42f46ab to
e5cdea3
Compare
Every other benchmark data generator sets zstd level 3 explicitly, since parquet-rs writes uncompressed by default. The three random-access generators were the remaining exception, so their files were uncompressed. Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xmiM7RPQBBN9ycUXegLDv
…workflow Add the action/bench-random-access-s3 label to the benchmarking guide's list of on-demand CI benchmarks, and drop a stray double blank line in the PR benchmark runner workflow. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L627uApatcD62FGrQPutCL Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Benchmarks: Random Access 📖Commits: PR How to read Verdict and Engines
vortex / arrow-ipc / ns (1.053x ➖, 0↑ 3↓)
random-access / vortex-file-compressed / ns (1.015x ➖, 0↑ 0↓)
random-access / parquet / ns (0.964x ➖, 2↑ 0↓)
random-access / lance / ns (1.010x ➖, 0↑ 0↓)
|
The Arrow IPC accessor only reads local disk, so the S3 variant was reporting local-disk timings under the arrow-tokio-s3 name. Refuse the format when a remote data dir is set and leave it out of the CI split. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L627uApatcD62FGrQPutCL Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
The three synthetic random-access generators (feature_vectors, nested_lists, nested_structs) wrote Parquet with default writer properties: uncompressed and, since every dataset is under the 1Mi-row default, one row group per file. Readers select row groups before rows, so a point lookup had to decode the whole file. Write them with 32Ki-row row groups, 1024-row data pages and zstd level 3, matching the other benchmark data generators. The row group size is a whole multiple of the 1024-row batches the Parquet-to-Vortex conversion reads, so the derived Vortex files are byte-identical. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L627uApatcD62FGrQPutCL Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
…-access-s3-benchmark-f9zi6a Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk> # Conflicts: # vortex-bench/src/random_access/mod.rs
…arquet Only the codec changes: zstd level 3 like the other generators. Row group and data page sizes stay at the parquet-rs defaults. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L627uApatcD62FGrQPutCL Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
…-access-s3-benchmark-f9zi6a Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
…ss-s3-benchmark-f9zi6a Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk> # Conflicts: # vortex-bench/src/random_access/mod.rs
v3 random-access records carry no storage field, and their measurement ID hashes only commit, dataset, format and open mode. The develop S3 run would therefore share IDs with, and overwrite, the local-disk rows. Give S3 runs an `-s3` dataset suffix (e.g. `taxi-s3/uniform`) so they form their own series without a schema change; local-disk IDs are unchanged. Also drop the cosmetic job name added to develop-bench.yml. Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L627uApatcD62FGrQPutCL Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
…ss-s3-benchmark-f9zi6a Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Summary
The random-access benchmark only ever read its data from local NVMe, which hides the cost that matters most for point lookups: how many bytes and round trips a format needs against an object store. This adds an S3 variant of the same benchmark so Vortex, Parquet, and Lance random access can be tracked against remote storage, both on
developand on demand for PRs.The synthetic Parquet files keep the parquet-rs default row group and page sizes; only the codec (zstd level 3, from #10157) is set.
Changes
random-access-benchgains--remote-data-dir s3://bucket/prefix/and--prepare-data. Data is materialized locally as before, uploaded verbatim, and then opened throughobject_store(Vortex and Parquet) or Lance's owns3://provider. Remote measurements are named...-tokio-s3and tagged withs3storage so they form a separate series from the local-disk numbers. Arrow IPC has no object-store reader and is skipped for remote runs.-s3dataset suffix (e.g.taxi-s3/uniform). This keeps them from overwriting the local-disk rows without a schema change; local-disk IDs are unchanged.vortex-benchaddsRemoteDataDir, which maps a local data path to its object key, plusopen_object_storeconstructors for the Vortex and Parquet accessors. The Parquet accessor keeps its cached footer and re-opens through a boxedAsyncFileReaderon each take, so the local and object-store paths share one read path.pr-bench-random-access-s3.ymlreusable workflow, wired intopr-bench-dispatch.ymlbehind theaction/bench-random-access-s3label andaction/bench-all.pr-bench-runner.ymllearnsvariant_idandremote_data_dirinputs, uploads the data before the run and deletes it afterwards.develop-bench.ymlgains aRandom Access (S3)matrix entry that refreshess3://vortex-ci-benchmark-datasets/develop/random-access/on every run.lance-benchenables Lance'sawsfeature soDataset::openacceptss3://URIs.Unit tests cover the key and URI mapping of
RemoteDataDirand the separate S3 ingest dataset.🤖 Generated with Claude Code
https://claude.ai/code/session_01L627uApatcD62FGrQPutCL