Skip to content

feat: add concurrent SNOS preparation and timing - #527

Open
heemankv wants to merge 18 commits into
mainfrom
feat/snos-log-instrumentation
Open

heemankv wants to merge 18 commits into
mainfrom
feat/snos-log-instrumentation

Conversation

@heemankv

@heemankv heemankv commented Jun 24, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • instrument retry-wrapped RPC calls with concurrency-safe per-run wait timing and method counts
  • split SNOS into async preparation and synchronous Cairo execution while preserving the compatibility API
  • make storage-proof batch size and in-flight request count configurable
  • process independent contract proofs concurrently
  • fetch current and previous block proofs concurrently
  • preserve deterministic key/chunk ordering so recorded witnesses replay across processes

Why

PC Devnet SNOS jobs spent roughly 91% of their wall time waiting on tens of thousands of RPC reads. The split API allows orchestrators to overlap RPC-bound preparation while independently controlling the memory-heavy Cairo finalization lane. Larger proof batches and bounded proof concurrency also remove most of the request amplification during upstream witness generation.

PC Devnet validation

Combined with Madara PR #1194's shared historical rollback overlays, the exact former 25m36.398s worst-case witness built in 61.982s on the same 16-vCPU node (24.8x). A ten-block sample averaged 58.818s and 50.757s after the cold first block. The resulting witness replayed successfully, validated its Cairo PIE, and spent 0.953s in measured replay/RPC wait.

Validated client settings:

SNOS_MAX_STORAGE_KEYS_PER_PROOF_REQUEST=5000
SNOS_MAX_CONCURRENT_PROOF_REQUESTS=4
SNOS_MAX_CONCURRENT_PROOF_CONTRACTS=7
SNOS_RPC_REQUEST_TIMEOUT_SECS=600

Consumer

Validation

  • GitHub Actions: rustfmt passed
  • GitHub Actions: Clippy passed
  • GitHub Actions: tests passed
  • cross-process replay of the optimized block-589720 witness passed Cairo PIE validation

@heemankv
heemankv requested review from Mohiiit and prkpndy as code owners June 24, 2026 07:02
@heemankv heemankv changed the title feat: add SNOS processing timing instrumentation feat: add concurrent SNOS preparation and timing Oct 5, 2026
@heemankv

heemankv commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Live PC Devnet validation of the deterministic RpcWitness path:

  • RPC control, block 584653: 1,675,153 ms total with 1,563,552 ms RPC wait and 30,602 calls.
  • Corrected live proactive witness, block 588384: 32,956,783-byte gzip, built upstream in 50.6 s.
  • Replay completed successfully in 116,262 ms: 1,484 ms fetch, 5,212 ms preparation, 702 ms measured replay/RPC wait, 109,565 ms OS execution, 22,385 local responses, and 27,297,869 Cairo steps.
  • Peak observed runner memory was about 1.8 GiB with ZIP disabled.

This confirms the sorted/deduplicated keys and deterministic chunk merge replay correctly across processes, and reduces per-job network wait from roughly 26 minutes to under one second.

@heemankv
heemankv force-pushed the feat/snos-log-instrumentation branch from e7a17fe to db8ccdd Compare October 6, 2026 13:58
@heemankv

heemankv commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Four-builder production-throughput result

Follow-up validation used the final PR images on the frozen PC Devnet database with four concurrent witness builds and 30 rollback-overlay cache entries.

  • contiguous witnesses: 310/310 (blocks 589800–590109)
  • wall time: 4,551.792s
  • effective throughput: 14.683s/witness, 4.086 witnesses/minute
  • throughput improvement versus the former 25m36.398s serial build: 104.6x
  • compressed artifact size: 46.0 MB average (33.1–64.7 MB)
  • no missing manifest, Kubernetes warning/OOM event, or producer restart

This narrowly beats the ~15.7s/block arrival interval measured in the earlier Mocknet soak. The production 16-vCPU/64-GiB configuration is therefore four builders with MADARA_RPC_HISTORICAL_OVERLAY_CACHE_ENTRIES=30.

An independently executed SNOS replay of block 589800 from the concurrent run succeeded and validated its Cairo PIE: 143.695s total, 0.842s witness fetch, 0.955s measured replay/RPC wait, 4.587s preparation, 138.265s Cairo execution, and 32,838,930 Cairo steps. This verifies that concurrent witness recording remains isolated and replay-correct.

@heemankv

heemankv commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Live sync plus witness-generation result\n\nThe optimized four-builder producer was run against the live Mocknet sequencer from a restored database.\n\n- exact first-hour completions: 241 witnesses\n- first-to-last completion span: 3,465s\n- effective creation rate: 14.438s/witness\n- response count: 29,043 average (19,868–37,000)\n- compressed size: 45.2 MB average (33.2–58.0 MB)\n- producer failures / OOMs / restarts: 0 / 0 / 0\n\nThat is faster than the approximately 15.7s Mocknet block interval measured in the earlier soak, while the same process continued syncing.\n\nThe long soak also exposed excessive per-proof INFO logging, so SNOS head c5e2a42 demotes contract/chunk/key details to DEBUG; Madara head c9b3957dc pins it. SNOS CI is green and the refreshed Madara CI is green through lint/build/orchestrator tests, with the final large test jobs still running.\n\nOperational note: the benchmark intentionally shared one PVC for RocksDB and witness output and became storage-I/O-bound later in deep catch-up. The production manifest uses a separate 1-TiB witness PVC.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant