Skip to content

Latest commit

 

History

History
693 lines (579 loc) · 46.3 KB

File metadata and controls

693 lines (579 loc) · 46.3 KB

Benchmark

The supported OpenVMM benchmark coordinator provides acceptance and diagnostic suites, a 21-metric microVM non-Python workload suite, and five device operation-rate metrics on Linux/KVM, Linux/MSHV, and Windows/WHP. At one vCPU, CI combines those workloads with eight 128 MiB shell lifecycle metrics and reports all 34 median (p50) values. At 2, 4, and 8 vCPUs, CI records only shell_snapshot_restore_512_mib. Latency and resident-memory metrics are lower-is-better; throughput and operation-rate metrics are higher-is-better.

The suite uses the base Alpine guest. Python application snapshots, the Python-agent console workload, and its snapshot-prefetch experiment are intentionally excluded because the supported guest build does not include a Python initramfs.

Human-readable timing summaries and lifecycle JSON report p50 and nearest-rank p95. The tracked platform CSVs continue to persist and gate p50 so the historical schema remains unchanged.

Lifecycle results report peak resident set size (RSS) for the measured OpenVMM process. Linux reads the process high-water mark; Windows reads the cumulative peak working set. RSS includes resident guest-memory mappings. CI persists and gates p50 RSS and reports both p50 and maximum RSS in its lifecycle diagnostics.

The coordinator samples peak RSS once, at the readiness marker, while OpenVMM is still running. A prequeued guest exit can end OpenVMM before that sample. The coordinator then discards and repeats the attempt, up to three attempts per measured sample, and reports the number of discarded attempts in peak_rss_remeasured_count. It never substitutes another reading: a sample taken before the marker omits the restore, Linux exit accounting includes the coordinator's pre-exec image, and a Windows reading after exit includes teardown.

Use this page for metric names and methodology. Current historical p50 values live in data/; timings copied into old discussions or commit messages are not baselines. Bare-metal and virtual-machine results have separate histories and must not be compared as one regression series. ABI version and processor count are also separate history dimensions. Legacy CSV rows are interpreted as ABI v1 with one vCPU; they are never used as an ABI-v2 one-vCPU baseline. New runs always emit ABI value 2 and remain stored under microvm-v2 paths so they cannot collide with legacy unsuffixed ABI-1 history. The OpenVMM benchmark coordinator is implemented in scripts/nvx_tools/benchmark.py and exposed through the supported NVX CLI. Guest payloads live in scripts/nvx_tools/benchmark_scripts; the coordinator fills the .sh.in templates before use.

CI performance series Backend Host type
linux-kvm-virtual-machine KVM Virtual machine
linux-mshv-virtual-machine MSHV Virtual machine
windows-whp-virtual-machine WHP Virtual machine

The series names carry no CPU vendor: every CI microVM runner is an Azure virtual machine with an Intel Xeon CPU (Ice Lake-SP or Emerald Rapids), so each series has one vendor's samples. No CI microVM runner has an AMD CPU, so no AMD series exists yet. An AMD runner needs a built-in AMD CPU profile, because CI never uses host profiles, and only Milan CPUs have one so far, amd.milan.v1. Such a runner boots its guests on that profile through AMD-V, so its timings would skew an Intel series' baseline. Before an AMD runner joins CI, give it series of its own, named with an -amd suffix such as linux-kvm-virtual-machine-amd, in PLATFORM_NAMES and OPENVMM_BACKENDS in scripts/nvx_tools/performance.py and in the CI matrix, so their history lives in separate data/ files. Local measurements on AMD hosts can use the names above, because only CI records history.

CI runs the complete acceptance and performance suites and the five device metrics at one vCPU under the canonical microVM, and only the 512 MiB shell snapshot restore at 2, 4, and 8 vCPUs. This produces 37 p50 values per series, 34 at one vCPU and one at each higher count, and 111 values across the three-series matrix. Counts run sequentially on each host so benchmark workloads never overlap on the same physical host.

Running locally

Build the release VMM and guest artifacts first. See Build.

Run the acceptance and diagnostic suites with:

python3 scripts/nvx.py benchmark --suite boot --backend whp
python3 scripts/nvx.py benchmark --suite e2e --backend kvm --platform linux-kvm-baremetal --processors 1 --memory-mib 128 --output data/runs/linux-kvm-baremetal/microvm-v2/1vcpu/acceptance.json

Run the complete performance suite with:

# Linux/KVM: run all 21 metrics and write collector-compatible logs
python3 scripts/nvx.py benchmark --suite performance --backend kvm --platform linux-kvm-baremetal --processors 1 --runs 5 --virtfs-runs 3 --skip-build --output-dir data/runs/linux-kvm-baremetal/microvm-v2/1vcpu

# Linux/MSHV: run all 21 metrics
python3 scripts/nvx.py benchmark --suite performance --backend mshv --platform linux-mshv-baremetal --processors 1 --runs 5 --virtfs-runs 3 --skip-build --output-dir data/runs/linux-mshv-baremetal/microvm-v2/1vcpu

# Windows/WHP
python scripts\nvx.py benchmark --suite performance --backend whp --platform windows-whp-baremetal --processors 1 --runs 5 --virtfs-runs 3 --skip-build --output-dir data\runs\windows-whp-baremetal\microvm-v2\1vcpu

Run the canonical device operation-rate suite with its default five warmups, 30 retained attempts per device, ten-second operation windows, and 512 MiB backing objects:

# Linux/KVM; use --backend mshv and the matching platform on Linux/MSHV.
python3 scripts/nvx.py benchmark --suite device-io --backend kvm --platform linux-kvm-baremetal --processors 1 --skip-build --output-dir data/runs/linux-kvm-baremetal/microvm-v2/1vcpu/device-io
python3 scripts/nvx.py performance collect --platform linux-kvm-baremetal --commit HEAD --input-dir data/runs/linux-kvm-baremetal/microvm-v2/1vcpu/device-io --output-dir data/results

# Windows/WHP.
python scripts\nvx.py benchmark --suite device-io --backend whp --platform windows-whp-baremetal --processors 1 --skip-build --output-dir data\runs\windows-whp-baremetal\microvm-v2\1vcpu\device-io
python scripts\nvx.py performance collect --platform windows-whp-baremetal --commit HEAD --input-dir data\runs\windows-whp-baremetal\microvm-v2\1vcpu\device-io --output-dir data\results

For a smoke test, pass --warmups 0 --runs 1 --device-io-duration-seconds 1.

Run the restore-only shape used by CI for higher-vCPU coverage with:

python3 scripts/nvx.py benchmark --suite shell-snapshot-restore --backend kvm --platform linux-kvm-baremetal --processors 8 --shell-memories 512 --warmups 1 --runs 5 --skip-build --output-dir data/runs/linux-kvm-baremetal/microvm-v2/8vcpu
python3 scripts/nvx.py performance collect --platform linux-kvm-baremetal --commit HEAD --input-dir data/runs/linux-kvm-baremetal/microvm-v2/8vcpu --output-dir data/results --require-shell-snapshot-restore-512

Measure restore-time vCPU activation from one immutable capacity-8 snapshot that boots with only CPU 0 online:

python3 scripts/nvx.py benchmark --suite snapshot-restore-vcpu --backend kvm --processors 8 --memory-mib 512 --warmups 1 --runs 5 --skip-build --output-dir data/runs/restore-vcpu-kvm

Use --backend mshv on Linux/MSHV or --backend whp on Windows/WHP. The suite captures once with maxcpus=1, then restores that same artifact with online targets 1, 2, 4, and 8. Each sample reaches its marker only after the guest has verified the requested online prefix. Capture and restore use --timeout (10 seconds by default). Results are diagnostic and are not fed into the historical fixed-vCPU performance CSVs.

For an explicit MSHV target, OpenVMM instantiates and binds only the requested VP prefix while retaining full-capacity topology and saved-state validation. MSHV restores without a target, and all KVM and WHP restores, instantiate the full capacity. A reduced-prefix MSHV runtime cannot be saved again.

Add --snapshot-profile to print p50 and p95 lifecycle phases for every online target. The diagnostic retains its pre-banner capture so one immutable boot-online-1 snapshot can serve every target. Its total restore latency is not comparable with the lifecycle-aligned shell-snapshot-restore metric; compare host-side lifecycle phases to separate VP binding and worker construction from guest resume behavior:

python3 scripts/nvx.py benchmark --suite snapshot-restore-vcpu --backend mshv --processors 8 --memory-mib 128 --warmups 1 --runs 5 --snapshot-profile --skip-build --output-dir data/runs/restore-vcpu-mshv-capacity-8

Measure restore-time memory activation from one immutable 512 MiB snapshot with 2 GiB of ABI-reserved capacity:

python3 scripts/nvx.py benchmark --suite snapshot-restore-memory --backend kvm --processors 1 --warmups 1 --runs 5 --snapshot-profile --skip-build --output-dir data/runs/restore-memory-kvm

Use --backend mshv on Linux/MSHV or --backend whp on Windows/WHP. The suite restores the same base snapshot at 512 MiB, 1 GiB, and 2 GiB. It reports guest-observed add-and-online latency separately from process-launch-to-ready latency, OpenVMM peak RSS, and optional host lifecycle phases. Expansion ranges are registered before restored execution; the guest marker is emitted only after every 128 MiB memory block is online.

The explicit 512 MiB target is the zero-expansion control. Its private status does not select the restore packet or request a snapshot boundary, so its launch-to-ready latency should remain within measurement noise of restoring a snapshot captured directly at 512 MiB.

Run the diagnostic snapshot lifecycle matrix with:

# Linux/KVM or Linux/MSHV
python3 scripts/nvx.py benchmark --suite snapshot-profile --backend kvm --warmups 1 --runs 5 --output data/runs/linux-kvm-baremetal/snapshot-profile.json

# Windows/WHP
python scripts\nvx.py benchmark --suite snapshot-profile --backend whp --warmups 1 --runs 5 --output data\runs\windows-whp-baremetal\snapshot-profile.json

The default matrix profiles 128, 256, 512, and 1024 MiB snapshots with both warm and cold restore artifacts. Use --shell-memories to select sizes and --cache-state warm, cold, or both to select cache conditions. The suite enables OpenVMM profiling for these diagnostic runs; other paths leave full profiling disabled unless --snapshot-profile is explicit. Snapshot capture always collects the OpenVMM clock records needed for its generation metric.

Run the non-canonical virtio restore diagnostic with a fresh output directory:

python3 scripts/nvx.py benchmark --suite device-restore-profile --backend kvm --platform linux-kvm-baremetal --processors 1 --warmups 1 --runs 5 --skip-build --output-dir data/runs/device-restore-kvm
python scripts\nvx.py benchmark --suite device-restore-profile --backend whp --platform windows-whp-baremetal --processors 1 --warmups 1 --runs 5 --skip-build --output-dir data\runs\device-restore-whp

By default the suite runs console, network, and virtio-fs in both active and deferred modes. Use --restore-devices and --restore-modes to select a subset. Deferred capture unbinds the selected built-in guest driver before nvx-snapshot; restore rebinds it and verifies I/O. Devices are located by the microVM ABI slots and virtio device IDs, not by unstable virtioN numbering: network at 0xd0000000/ID 1, virtio-fs at 0xd0001000/ID 26, and console at 0xd0002000/ID 3. Probe control markers use port-B (/dev/hvc0) independently of the selected device. Before rebinding, deferred mode sends a descriptor-free queue-0 MMIO notification through /sbin/nvx-mmio-write to verify staged-kick ordering. Console verification writes through /dev/hvc1; network uses the portable profile and a gateway ping; virtio-fs reads a host seed and writes a host-visible result after remount.

The output directory contains device-restore-profile.json, benchmark-metadata.json, and raw capture/restore logs. The JSON retains every measured sample, p50/p95 launch-to-ready and trigger-to-first-I/O timing, peak RSS, guest markers, structured restore event order, queue-start counts, staged-kick dispatch counts, and the required zero stale/premature callback assertion. Only this suite enables the virtio_restore=debug event stream. Its result is diagnostic: it is not accepted by performance collect, persisted to data/*.csv, or used by performance gate. Host-observed marker timing includes serial delivery and scheduler delay, and repeated restores reuse one fresh snapshot per scenario; compare results only on the same host under equivalent load and power conditions.

Run one workload by selecting cold-start, device-io, virtfs, shell-snapshot, or network-snapshot instead of performance. Use --shell-memories 128 256 512, --payload-mib 64, and --net 10.0.0.2/24 --network-profile portable to override their defaults. Run python scripts/nvx.py benchmark --help for the complete option surface.

The network benchmarks select the same in-process portable data plane on KVM, MSHV, and WHP. They do not create TAP devices or require host firewall rules.

Each workload directory includes benchmark-metadata.json with the platform, backend, ABI, processor count, host affinity set, memory sizes, artifact revisions, repository dirty state, warmups, measured run counts, and SHA-256 values for the VMM, kernel, initramfs, and coordinator. Device metadata also records helper source and packaged-binary hashes. Collection rejects mismatched lifecycle/workload metadata and duplicate topology rows.

Use a fixed affinity set containing one logical processor per physical core and at least N+2 processors for an N-vCPU guest. The additional processors cover VMM and device work. The default selector follows this policy; an explicit undersized --cpus set is rejected before measurement.

The virtual-machine CI runners expose four cores as eight sibling logical CPUs. Their virtual-machine series deliberately use the fixed 0-7 set with --host-cpu-reserve 0, including sibling CPUs and sharing capacity between guest, VMM, and device work. These constrained nested-host results are kept separate from the bare-metal series; other runs retain the two-CPU reserve. CI gives the constrained series a 40-second guest-marker deadline; completed measurements still record their actual latency.

MSHV lifecycle diagnostics

The Linux/MSHV backend emits an opt-in MSHV_SET_GUEST_MEMORY completed event from the existing mshv map user memory span. The span identifies the guest range and permissions; the event records elapsed_us and success. On x86_64, MSHV_CREATE_VCPU completed reports BSP creation after RAM attachment. The backend-neutral post-memory partition finalization completed event includes BSP creation and capability discovery. These scopes are nested and must not be summed. Enable virt_mshv=info,openvmm_core::worker::dispatch=info only for diagnostic run invocations. Benchmark measurements force OpenVMM logging off to avoid changing the measured path.

Collect a completed suite with:

python3 scripts/nvx.py performance collect --platform linux-kvm-baremetal --commit HEAD --input-dir data/runs/linux-kvm-baremetal/microvm-v2/8vcpu --output-dir data/results --require-network --require-shell-snapshot --require-shared-suite --lifecycle-input data/runs/linux-kvm-baremetal/microvm-v2/8vcpu/acceptance.json

Benchmark commands

Benchmark Command Description
Shell lifecycle benchmark --suite e2e Measures cold start, snapshot generation, snapshot restore, teardown, and peak RSS using a shell-ready guest.
All supported non-Python workloads benchmark --suite performance Runs 21 metrics and writes collector-compatible logs.
Device operation rates benchmark --suite device-io Measures five random storage and UDP round-trip operation rates with resumable raw attempts.
Cold start benchmark --suite cold-start Measures a quiet shell-ready baseline and isolated one-parameter kernel command-line variants.
Virtual file system benchmark --suite virtfs Measures live host-directory throughput and verifies host-to-guest plus guest-to-host visibility in one running VM.
Shell snapshot benchmark --suite shell-snapshot Compares cold boot with lifecycle-aligned snapshot restore at 128, 256, and 512 MiB.
Shell snapshot restore benchmark --suite shell-snapshot-restore Captures an unmeasured lifecycle-aligned snapshot and measures only restore latency for the selected memory sizes.
Restore-time vCPU activation benchmark --suite snapshot-restore-vcpu --processors 8 Restores one boot-online-1, capacity-8 snapshot at online targets 1/2/4/8 and reports latency plus peak RSS.
Network snapshot benchmark --suite network-snapshot Compares a network-ready cold boot with snapshot restore and verifies gateway connectivity.
Snapshot lifecycle profile benchmark --suite snapshot-profile Retains raw capture and restore phase records and summarizes 128/256/512/1024 MiB warm/cold restores.
Virtio device restore profile benchmark --suite device-restore-profile Verifies active and driver-unbound deferred restore for console, network, and virtio-fs; emits standalone diagnostic JSON and raw logs.

Kernel command lines

Each benchmark boots the guest with a fixed kernel command line. Two bases recur below:

  • BASE = earlycon=xe9 console=hvc0 reboot=t panic=-1
  • QUIET = earlycon=xe9 console=hvc0 quiet loglevel=0 reboot=t panic=-1

OpenVMM owns BASE; each --cmdline value below is appended to it. Restore phases do not pass a command line and resume the one captured in the snapshot. OpenVMM also appends device-discovery tokens (virtio_mmio.device=... for an attached mount or NIC) and may replace console=hvc0 with console=hvc1 when a virtio console is selected. It appends no clock token: under the time ABI the NVX kernel reads the TSC and LAPIC timer rates from the hypervisor identity's frequency MSRs and never calibrates either clock, and OpenVMM rejects a supplied tsc_early_khz= or lapic_timer_hz= token (E_CMDLINE_CLOCK_TOKEN). No benchmark command line tunes the clock or clocksource, except the cold_start_clocksource variant below.

Benchmark Phase Kernel command line
Shell lifecycle cold/capture BASE plus BASE_TUNING: random.trust_cpu=on rcupdate.rcu_expedited=1 nokaslr mitigations=off cryptomgr.notests quiet loglevel=0; CI uses 128 MiB.
Shell lifecycle restore restore (from the measured shell-ready snapshot)
Cold start baseline QUIET
Cold start tuning variant QUIET plus one of clocksource=tsc, tsc=reliable, no_timer_check, random.trust_cpu=on, rcupdate.rcu_expedited=1, nokaslr, mitigations=off, or cryptomgr.notests
Virtual file system guest runs QUIET
Device operation rates guest runs QUIET; virtio-net uses the portable profile at 10.0.0.2/24
Shell snapshot cold QUIET
Shell snapshot capture QUIET; host-driven after the boot marker and SMP/LAPIC probe
Shell snapshot restore restore through the post-restore SMP/LAPIC probe and lifecycle restore marker
Network snapshot cold QUIET virtnet_probe=<gateway>
Network snapshot capture QUIET virtnet_probe=<gateway> netsnap
Network snapshot restore restore (from snapshot)

Canonical metrics

Shell lifecycle

CI runs one warmup and ten measured samples with a 128 MiB guest. Each phase uses a fresh OpenVMM process. The final measured snapshot is retained for the restore samples.

Metric Unit Description
openvmm_cold_start ms Immediately before OpenVMM process creation through ALPINE-MICROVM-BOOT-OK.
openvmm_snapshot_generation ms OpenVMM process-clock interval from capture.input_gate start through capture.publication_commit; excludes host-to-guest console delivery and publication polling.
openvmm_snapshot_restore ms Immediately before restored OpenVMM process creation through OPENVMM-SNAPSHOT-RESTORE-OK.
openvmm_cold_start_guest_exit_teardown ms Dispatch of guest nvx-exit 0 after the cold-start marker through successful OpenVMM process exit.
openvmm_snapshot_restore_guest_exit_teardown ms Host observation of the restore marker through successful OpenVMM process exit, with guest nvx-exit 0 already queued in the capture controller.
openvmm_cold_start_peak_rss MiB Per-process peak RSS through the cold-start marker.
openvmm_snapshot_generation_peak_rss MiB Per-process peak RSS for the snapshot-generating process.
openvmm_snapshot_restore_peak_rss MiB Per-process peak RSS through the restore marker.

Warmups are excluded from every aggregate. The CSV stores and gates p50 values. CI retains the per-phase lifecycle profile, and diagnostics additionally report timing minimum, maximum, and sample count plus peak-RSS maximum. Collection rejects host-termination semantics, legacy console-timed capture results, restore results without prequeued guest exit, missing samples, and any guest-exit teardown timeout. With at least ten samples, it also rejects snapshot-generation series whose p50 is more than 25% above p25; this prevents a transient host stall from entering performance history without hiding a uniformly slower product result. Every launch is scanned for the guest's time ABI output as in the microVM correctness jobs: a violation event or a time ABI power-off fails the run. Each teardown that can return a sample (a guest exit with status 0, a host termination, or a teardown timeout) first reads the console to its end, so an event that the guest prints after the readiness marker, even on an unterminated last line, still fails the run. A host-terminated launch, or one killed after a teardown timeout, fails on a time ABI power-off status but not on the status of its termination. A guest exit with a nonzero status fails the run regardless; the scan reads what the console delivers within a second, so that a time ABI power-off is reported with its event. Benchmarks never ask the guest for its time ABI status, as the correctness jobs do after cold boots and restores, so the guest prints nothing extra. Up to the marker, the scan only parses output the harness reads anyway, and the harness reads the rest of the console only after it records a sample's intervals, so the scan adds nothing to them. The guest's boot check, which powers the guest off with status 193 if it fails, is part of openvmm_cold_start. Lifecycle capture runs a deterministic affinity-pinned worker on every vCPU before the snapshot request. Explicit correctness scenarios also stage a post-restore probe. The capture probe is outside the snapshot-generation timing interval. Each worker proves that it executed on its assigned vCPU and observes that CPU's local APIC counter advance. Workers poll for at most 10,000 counter reads, so a stalled timer fails without relying on guest sleeps or a working guest clock to bound the check. Pre/post interrupt snapshots also require every counter to advance and remain at least as large as the worker's observed value. The check remains strict even on a one-vCPU guest: falling back to the PIT after a failed LAPIC calibration is not success. A counter frozen at 12 can indicate Linux's APIC timer disabled due to verification failure; increasing the poll budget cannot repair it. The time ABI hides the TSC-deadline timer on every backend, so every guest uses the one-shot counting LAPIC, and the smp scenario that the microVM correctness jobs run at 1, 2, 4, and 8 vCPUs covers it. The benchmarks run no LAPIC gate of their own (#286), and no guest boots with lapic=notscdeadline. test-microvm --scenario smp-lapic remains for explicit local use: it runs smp and also asserts that no CPU lists tsc_deadline_timer, that each online CPU's clockevent device is the counting lapic, and that the boot line reports the backend's LAPIC rate for every CPU. No default suite or CI job runs it, so nothing runs twice. Captures do not wait for a clocksource: under the time ABI, Linux registers the tsc clocksource at device_initcall, before any guest work, so no capture can observe the transitional tsc-early window. The coordinator stages the probe and a capture controller in guest memory. The controller runs the first probe, blocks in read, and invokes nvx-snapshot when the host sends the trigger. The controller always emits NVX-SNAPSHOT-DISPATCHED immediately before nvx-snapshot. On restore it completes any staged validation, prints the restore marker, and executes the prequeued nvx-exit 0 in guest-exit mode. Host-termination mode leaves the guest running until the host terminates it.

The interrupt-less hvc0 console polls even when the controller blocks in read. Its delivery delay is retained in the non-gating request_to_publication_* JSON fields; a profiled capture also separates console_command_round_trip and host-observed guest_dispatch_to_publication. Neither observer interval is the generation metric. The canonical samples_ms/p50_ms and profile phase capture.snapshot_generation use the same OpenVMM clock interval.

Cold start

These nine scenarios normally use five samples. All measure milliseconds from OpenVMM process launch to the shell-ready console marker at the configured memory size. The tuning scenarios append exactly one kernel parameter to the quiet baseline.

Metric Description
cold_start_base Quiet baseline with no additional tuning parameter.
cold_start_clocksource Baseline plus clocksource=tsc on every backend. Before the time ABI, KVM measured clocksource=kvm-clock, so KVM history before that change measures a different clocksource.
cold_start_tsc_reliable Baseline plus tsc=reliable.
cold_start_no_timer_check Baseline plus no_timer_check.
cold_start_random_trust_cpu Baseline plus random.trust_cpu=on.
cold_start_rcu_expedited Baseline plus rcupdate.rcu_expedited=1.
cold_start_nokaslr Baseline plus nokaslr.
cold_start_mitigations_off Baseline plus mitigations=off.
cold_start_cryptomgr_notests Baseline plus cryptomgr.notests.

Virtual file system

Throughput scenarios normally use three samples and a 64 MiB payload. The round-trip scenario measures complete process wall time while host and guest exchange files through one running VM.

Metric Unit Description
virtfs_live_write MB/s Sequential guest dd write with fsync directly into the host directory.
virtfs_live_read MB/s Sequential guest read from the host file after dropping guest page cache.
virtfs_live_roundtrip ms Guest creates a marker observed by the host, then observes a host rewrite before that VM exits.

Device operation rates

The dedicated suite always uses the canonical microVM and attaches the block backing object as the writable scratch role. There is no ABI selector or legacy unroled --virtio-blk control. Every guest has 256 MiB RAM and runs the same static, dependency-free x86-64 helper from the reproducible initramfs. Canonical runs execute five warmups followed by 30 retained attempts for each device. Every operation window lasts ten seconds against a 512 MiB backing object.

Metric Unit Timed work
virtio_blk_random_read_iops ops/s QD1 aligned 4 KiB random pread64 with O_DIRECT.
virtio_blk_random_write_iops ops/s QD1 aligned 4 KiB random pwrite64 with O_DIRECT.
virtio_fs_random_read_iops ops/s QD1 buffered 4 KiB random pread64 after sync and a guest page-cache drop.
virtio_fs_random_write_iops ops/s QD1 buffered 4 KiB random pwrite64; the final flush starts after timing ends.
virtio_net_udp_roundtrip_ops ops/s Completed 64-byte UDP echo request/response transactions through portable virtio-net.

Storage offsets use the same deterministic pseudo-random sequence in every backend. The host UDP server binds the host's primary IPv4 address, requires no TAP or firewall changes, and stops before the suite returns. The rate is derived during aggregation as $\mathrm{ops/s}=N\times10^9/t_{ns}$ from integer operation counts and monotonic elapsed nanoseconds.

Backing objects are reused across attempts. On Windows, extending the virtio-fs file sets its length without initializing its contents, so the first random writes can include host filesystem initialization and zero-fill costs. A discarded warmup keeps those cold-file costs out of the steady-state operation rates; a single cold attempt is only a smoke test.

device-io.log stores one versioned JSON record per warmup or retained attempt. Missing, duplicate, malformed, or zero-work helper output turns that attempt into an explicit failure; failed attempts stay in the log and never enter p50 or p95. Warmup and retained indices are fixed, so a failure cannot shift sampling. Re-running with identical metadata skips completed identities and resumes the first absent attempt. The selected ABI is part of metadata and the default output directory, so v1 and v2 records cannot mix. Changed controls or provenance are rejected. Temporary raw disks, shared directories, the UDP server, and each benchmark-owned VMM are cleaned on success, failure, or interruption.

Shell snapshot

Each memory size normally uses five cold boots and five restores. Cold values run through the shell-ready ALPINE-MICROVM-BOOT-OK marker. The unmeasured capture then runs the same SMP/LAPIC probe as the shell lifecycle benchmark before requesting the snapshot. Restore values run through that probe and the standalone OPENVMM-SNAPSHOT-RESTORE-OK marker. Capture and restore boundaries therefore match the shell lifecycle benchmark; the kernel command lines remain different so the cold metrics are not interchangeable.

This lifecycle-aligned methodology supersedes the earlier pre-banner shellsnap capture. Historical shell_snapshot_restore_* values produced by that methodology are not comparable with newly collected values.

The 64 MiB pair, shell_snapshot_cold_64_mib and shell_snapshot_restore_64_mib, was retired in #116. The platform CSVs in data/ keep its results only as history; the last ones are b8912df's.

Metric Description
shell_snapshot_cold_128_mib OpenVMM launch to a shell-ready guest with 128 MiB of memory.
shell_snapshot_restore_128_mib Restore process launch through lifecycle-aligned verification of a 128 MiB snapshot.
shell_snapshot_cold_256_mib OpenVMM launch to a shell-ready guest with 256 MiB of memory.
shell_snapshot_restore_256_mib Restore process launch through lifecycle-aligned verification of a 256 MiB snapshot.
shell_snapshot_cold_512_mib OpenVMM launch to a shell-ready guest with 512 MiB of memory.
shell_snapshot_restore_512_mib Restore process launch through lifecycle-aligned verification of a 512 MiB snapshot.

Snapshot lifecycle profile

This diagnostic suite captures a fresh snapshot at each selected memory size, then restores that artifact under each selected cache condition. A warm run sequentially reads manifest.bin, state.bin, and memory.bin before every restore. Linux cold runs apply POSIX_FADV_DONTNEED to each artifact. Windows cold runs restore from a fresh unbuffered robocopy /J clone. The selected mechanism is recorded as cache_control alongside each result.

Snapshot capture always enables OPENVMM_STARTUP_PROFILE=1 to measure generation on OpenVMM's process-relative clock. Missing, duplicate, wrong-process, or non-monotonic capture boundaries are errors, not a fallback to console timing. Full profile retention and host resource counters remain opt-in through --snapshot-profile or the snapshot-profile suite. Other benchmark subprocesses remove the variable unless profiling is requested. OpenVMM writes one ASCII record per phase to stderr with the versioned OPENVMM_SNAPSHOT_PROFILE_V1 prefix. Each record contains:

  • operation, phase, and exclusive, where exclusive phases are disjoint intervals and non-exclusive phases are cumulative milestones that may contain nested work;
  • monotonic duration_ns and process-relative process_elapsed_ns, plus the emitting pid;
  • phase-specific logical_bytes, allocated_bytes, gpa_faults, and populated_bytes when available.

OpenVMM relays the guest console to stdout from its own thread, with no ordering against its stderr writes, so either stream can split a line of the other when they share a terminal or pipe. Snapshot capture and launch measurements therefore read stderr through a separate pipe: they parse profile records only from stderr and match guest markers only on the console. Their logs and error reports interleave the two streams by whole lines. Because the streams are read independently, a record written before a guest marker can arrive after it. Snapshot capture, and launch measurements before they return a sample, read both streams to their end within the configured timeout and fail if either stream does not end.

With full profiling, the coordinator retains every record in profile.raw_samples. It derives capture.snapshot_generation from the OpenVMM clock and adds observer-defined process_startup, console_input_dispatch, console_command_round_trip, guest_dispatch_to_publication, request_to_publication, source_teardown, resume_to_readiness, and process_launch_to_readiness boundaries without mixing them into OpenVMM-exclusive intervals. console_input_dispatch measures the synchronous write of the snapshot command and prequeued restore script to the OpenVMM console and records the payload size. All captures emit a guest marker immediately before nvx-snapshot; console_command_round_trip ends when the host observes that marker, and guest_dispatch_to_publication spans that observation through snapshot publication. Each observed record also includes available process counters. Linux reports RSS, peak RSS, minor and major faults, total page faults, and, when smaps_rollup is available, private dirty and private RSS bytes. Windows reports working set, peak working set, private commit, and page faults.

profile.phases groups samples by operation.phase and stores raw samples_ms plus p50, nearest-rank p95, minimum, and maximum. The top-level JSON path is snapshot_profile_matrix.<backend>.<memory_mib>. Each memory entry contains the capture result and restore.warm and/or restore.cold, including its cache control and full restore result. Warmups are excluded from raw samples and summaries.

Capture records isolate guest quiesce, state save, mapped-memory and memory-handle flushes, each publication step, publication observation, and source teardown. Restore records isolate artifact open and preparation, COW section and mapping/view creation, prototype and final partition work, GPA registration, VP-thread binding, saved-state restore, device start, generation-ID creation, the cumulative gated guest-repair interval, and guest resume to readiness. State and memory SHA stages are intentionally absent from the final capture and restore path.

startup.vp_thread_bind is the exclusive wall interval for all VP threads to bind. Nested startup.vp_bind_bsp and startup.vp_bind_ap_<INDEX> records are non-exclusive per-VP intervals; compare their endpoints with the aggregate interval to expose serialized backend VP creation.

Network snapshot

These scenarios normally use five samples. A run is accepted only after the network verification marker is observed. Cold boot and restore both time to one successful ICMP echo to the configured gateway; the one-second timeout bounds a failed probe without adding an interval between successful packets.

Metric Description
network_snapshot_cold OpenVMM launch to the network-ready marker after interface configuration and a real connectivity probe.
network_snapshot_restore Restore process launch to the verified restored-network marker.
network_snapshot_restore_wall End-to-end host process wall time for restoring the snapshot, rebuilding the network backend, verifying connectivity, and exiting.

CI collection

python scripts/nvx.py performance collect --require-shared-suite rejects a workload result unless it contains exactly the 21 shared metrics. Supplying --lifecycle-input requires and merges the eight lifecycle metrics, producing a 29-metric one-vCPU result. CI collects the five device operation-rate metrics from their separate raw-log directory and merges them by ABI, processor count, commit, and metric, producing the final 34-metric ABI-2 one-vCPU result. Higher-vCPU collection uses --require-shell-snapshot-restore-512, which requires exactly shell_snapshot_restore_512_mib plus canonical one-warmup/five-sample metadata for a 2-, 4-, or 8-vCPU guest. A device-io directory is recognized from metadata and must contain exactly its five metrics and every configured attempt. Each backend job publishes its p50 tables and one-vCPU lifecycle diagnostics to $GITHUB_STEP_SUMMARY. Pull-request regression checks compare KVM, MSHV, and WHP results with the latest base-branch history.

These jobs consume, gate, publish, or persist the KVM benchmark results (#286); MSHV and WHP follow the same flow with their own platform jobs:

  • Platform / Linux / KVM / Virtual machine (platform-kvm) runs the benchmarks and publishes the results: the run-scoped benchmark-linux-kvm-virtual-machine-<run-id> artifact, the per-attempt diagnostics artifact, and the p50 tables in its step summary.
  • Performance regression gate (performance-gate, pull requests only) consumes the three platform artifacts once all three platform jobs succeeded and gates them against the base branch's history. It publishes and persists nothing.
  • Persist performance baseline (performance-persist, dev pushes only) consumes them and persists them to data/*.csv.
  • Publish development release (release, dev pushes only) doesn't read them, but it and performance-persist run only when the three platform jobs succeeded and every correctness job, nvx-microvm-tests-kvm included, succeeded or was skipped.
  • Required status check fails unless every job that the change schedules, including platform-kvm, performance-gate, and nvx-microvm-tests-kvm, has its expected result.

Linux CI uses a reduced device contract of zero warmups and one retained attempt. Windows CI discards one warmup and retains five attempts so file initialization cannot dominate its p50. Both CI contracts use one-second operation windows. Canonical baseline collection uses the full 5 + 30 contract.

The current workflow collects 10 measured lifecycle samples after one warmup. Linux and Windows CI validate the lifecycle result before starting the remaining benchmark suites. Snapshot-generation instability is reported with temporary-failure exit status 75; CI remeasures it once on the same runner. Each attempt is preserved in the benchmark artifact as acceptance-attempt-1.json or acceptance-attempt-2.json, including its lifecycle profiles. Only a validated attempt is copied to acceptance.json for collection; a stale accepted result is removed before measuring. Other validation failures stop immediately, and a second unstable result still fails the job. The stability guard rejects a p50 more than 25% above p25 and also rejects an adjacent gap above 25% when at least two samples lie on each side. Singleton outliers remain tolerated, while pooled runners cannot publish a bimodal host-stall series into topology-wide history. On the Linux CI runners, snapshot generation takes about 1 ms on KVM and about 4 ms on MSHV, so a sub-millisecond CPU stall in two samples is enough to cross the gap limit.

Before updating the run-scoped benchmark-<platform>-<run-id> artifact, each platform uploads its raw results and lifecycle profiles to the immutable benchmark-diagnostics-<platform>-<run-id>-attempt-<run-attempt> artifact, including after a benchmark failure. Both artifacts are retained for one day. Rerunning a workflow cannot overwrite an earlier attempt's diagnostics. Downstream gates still use the run-scoped artifact so a failed-jobs-only rerun can reuse successful platform results from an earlier workflow attempt. The acceptance-attempt-N.json files identify the bounded lifecycle remeasurements within one workflow attempt, not the workflow's github.run_attempt.

When investigating instability, compare each attempt's snapshot_capture.<backend>.profile.raw_samples with its samples_ms. For example, a slow capture.mapped_memory_flush with otherwise stable capture phases localizes the delay to host-side mapped RAM flushing, not guest boot or snapshot restore. Extra time in capture.quiesce, capture.save_state, or between profiled phases while the flush and publication phases stay flat instead points to a host CPU stall. Reproduce with the exact executable and guest artifact hashes on the same host before attributing the delay to a source change. Keep the stability thresholds and bounded remeasurement unchanged when collecting diagnostic evidence.

Snapshot generation is bounded by the write throughput of the benchmark's temporary directory, because capture flushes guest RAM through a backing file created there. The Windows backing file stays dense so restore keeps the captured pages cached; NTFS therefore zero-fills the unwritten range below the guest's top-of-RAM pages, and a 128 MiB guest writes about 150 MiB per capture. The Windows CI runners' 128 GiB Premium SSD system disk also holds the runner work tree, caches, and builds. Under load it alternates between its burst limit of about 173 MB/s and its baseline of about 102 MB/s in blocks of tens of seconds, which splits one lifecycle series into flush clusters near 0.95 and 1.6 seconds. Windows CI therefore passes --scratch-dir with a per-job directory under the NVX_BENCHMARK_SCRATCH root that runner provisioning creates on the data volume. Acceptance JSON records the directory as controls.scratch_directory, and workload metadata records it as scratch_directory. When the root is not provisioned, CI warns and uses the system temporary directory. Windows-coordinated KVM workers receive the WSL translation of the same directory instead of falling back to WSL's /tmp. Use --scratch-dir for manual runs whose temporary directory shares a volume with other I/O-heavy work.

The regression gate compares the target p50 with the median of the latest 10 p50 values on the pull request's base branch and requires all 10 matching history points. A metric regresses only when it is more than 50% worse. Lower-is-better millisecond metrics must also be more than 10 ms slower; higher-is-better metrics use the percentage comparison alone. Missing or insufficient history is a warmup, not a failure. Successful dev builds append collected results to topology-specific files in data/. Every new ABI/count series begins as a warmup baseline before its regression gate has enough matching history.

Only the virtual-machine topologies exercised by CI have tracked histories. CI passes --history-reset-dir data to honor metrics explicitly removed from a tracked candidate history file even before the reset merges into the base branch. After merge, reset metrics start the normal ten-point warmup again; a missing candidate file alone does not reset a base-branch history.

Lifecycle methodology

The lifecycle benchmark uses the optimized 128 MiB shell-ready guest as one baseline across KVM, MSHV, and WHP. Host timing starts immediately before Popen, so cold-start and restore values include OpenVMM process startup and VM construction. Snapshot generation uses OpenVMM's monotonic process clock from the beginning of input gating, after the guest requests the snapshot, through atomic publication. It does not include delivery of nvx-snapshot over the polling console. Host request-to-publication and first publication observation (polled every 1 ms) remain diagnostics.

After a cold-start marker, the host dispatches guest nvx-exit 0 and measures until the OpenVMM process exits successfully; cold-start teardown still includes console delivery. Snapshot capture instead queues the restore marker and nvx-exit 0 together after the capture boundary. Restore teardown therefore begins at the standalone marker line and does not depend on a second host-to-guest poll of the interrupt-less hvc0 console. It is a host-observed marker-to-exit interval, including remaining guest shutdown and observer scheduling, not a measurement of VMM-internal teardown alone. CI allows up to 15 seconds for process exit before rejecting a measured sample. Snapshot-source exit after publication is retained in the raw JSON as a diagnostic but is not the guest-exit teardown metric.

Cold-start methodology

Cold-start timing begins immediately before launching the OpenVMM process and ends when the shell-ready console substring appears. It therefore includes host process startup, VM construction, and guest execution. Every scenario suppresses routine kernel logging and captures console output without rendering it, so terminal rendering is not part of the measurement.

The earlycon=xe9 and hvc0 paths use one outb per byte. They replace a 16550-style path that would normally require a line-status read plus a data write. The VMM scans each emitted byte synchronously for markers. OPENVMM_LOG=off suppresses VMM logs, while quiet loglevel=0 suppresses most output in the guest.

The baseline command line is:

earlycon=xe9 console=hvc0 quiet loglevel=0 reboot=t panic=-1

Each other scenario appends only the parameter named in the metric table. This isolates that parameter's effect; the benchmark does not combine the tunings or change guest memory between scenarios. Security-reducing parameters such as nokaslr and mitigations=off are measured for trusted, single-tenant deployments and are not general recommendations.

Major cold-start contributors identified while building the current kernel configuration were:

  • PM_TRACE_RTC wall-clock probing, which can wait on absent or minimal RTC behavior;
  • initialization and probing for hardware subsystems the machine does not expose;
  • per-page kernel metadata initialization as guest RAM grows;
  • console VM exits.

The checked-in kernel therefore omits unused PCI, storage, graphics, sound, power-management, tracing, and debug stacks, keeps a minimal RTC implementation, and uses a low tick rate with tickless idle. Treat these as design constraints when changing kernel/config-microvm; validate both boot correctness and the relevant cold-start metrics.