Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions .github/workflows/copilot-setup-steps.yml
Original file line number Diff line number Diff line change
Expand Up @@ -350,13 +350,26 @@ jobs:
if [[ -r /dev/kvm && -w /dev/kvm ]]; then
kvm=yes
fi
# The CPU as NVX names it (vendor family/model/stepping), which
# decides the CPU profile that a cold boot's --cpu-profile auto
# selects (doc/usage.md, "CPU profiles").
cpu=$(awk -F': *' '
/^vendor_id/ { vendor = $2 }
/^cpu family/ { family = $2 }
/^model[ \t]*:/ { model = $2 }
/^stepping/ { stepping = $2 }
/^model name/ { name = $2 }
/^$/ { exit }
END { printf "%s %s/%s/%s (%s)", vendor, family, model, stepping, name }
' /proc/cpuinfo)
{
echo "### NVX Copilot environment"
echo
echo "| Item | Value |"
echo "| --- | --- |"
echo "| Setup status | ${status} |"
echo "| Failed steps | ${failed_list:-none} |"
echo "| CPU | ${cpu} |"
echo "| CPUs / memory | $(nproc) / $(free -g | awk '/^Mem:/ { print $2 }') GiB |"
echo "| Workspace free space | $(df -h --output=avail . | tail -n 1 | tr -d ' ') |"
echo "| KVM access | ${kvm} |"
Expand Down
11 changes: 11 additions & 0 deletions doc/benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,17 @@ coordinator fills the `.sh.in` templates before use.
| `linux-mshv-virtual-machine` | MSHV | Virtual machine |
| `windows-whp-virtual-machine` | WHP | Virtual machine |

The series names carry no CPU vendor: every CI microVM runner is an Azure virtual machine with an
Intel Xeon CPU (Ice Lake-SP or Emerald Rapids), so each series has one vendor's samples. No CI
microVM runner has an AMD CPU, so no AMD series exists yet. An AMD runner needs a built-in AMD
[CPU profile](usage.md#cpu-profiles), because CI never uses host profiles, and only Milan CPUs have
one so far, `amd.milan.v1`. Such a runner boots its guests on that profile through AMD-V, so its
timings would skew an Intel series' baseline. Before an AMD runner joins CI, give it series of its
own, named with an `-amd` suffix such as `linux-kvm-virtual-machine-amd`, in `PLATFORM_NAMES` and
`OPENVMM_BACKENDS` in `scripts/nvx_tools/performance.py` and in the CI matrix, so their history
lives in separate `data/` files. Local measurements on AMD hosts can use the names above, because
only CI records history.

CI runs the complete acceptance and performance suites and the five device metrics at one vCPU
under the canonical microVM, and only the 512 MiB shell snapshot restore at `2`,
`4`, and `8` vCPUs. This produces 37 p50 values per series, 34 at one vCPU and one at each higher
Expand Down
2 changes: 1 addition & 1 deletion doc/ci.md
Original file line number Diff line number Diff line change
Expand Up @@ -337,7 +337,7 @@ microcode and the invariant-TSC flags, may read `unknown`.
| Check | Implementation |
| --- | --- |
| H1 | `/dev/kvm` or `/dev/mshv` is readable and writable, and a KVM host has no `/dev/mshv`; on Windows, `WHvGetCapability` reports a hypervisor |
| H2 | Vendor, family, model, stepping, microcode, and OS build from `/proc/cpuinfo` or the Windows registry. Then `openvmm --hypervisor <backend> --cpu-fingerprint <path>` writes the host's CPU fingerprint (by default `nvx-cpu-fingerprint-<backend>.json` in the probe directory; `--cpu-fingerprint` overrides it), checks it against the profile that `auto` selects from OpenVMM's catalog, which shares one profile per generation across backends, and prints one `NVX-CPU-PROFILE:` line. H2 reports the generation and the profile from that line, so a profile that an OpenVMM pin adds qualifies its hosts without an NVX change. H2 requires exit status 0, `status=pass`, the same backend, and a catalog profile of the reported generation, reports the profile and surface digests, and otherwise fails with OpenVMM's code, for example `[E_PROFILE_HOST_UNKNOWN]` on a CPU that no profile serves or `[E_PROFILE_UNSUPPORTED]` naming every unsupported CPUID bit. On WHP, OpenVMM checks the entries outside the profile on a probe partition configured from the profile, as a cold boot does. With `--no-openvmm`, H2 maps the host with NVX's copy of the catalog instead: `skylake-sp` (6/85, steppings 0 to 4, `intel.skylake-sp.v1`), `icelake-sp` (6/106, `intel.icelake-sp.v1`), `emeraldrapids` (6/207, `intel.emeraldrapids.v1`), or `alderlake` (6/151 and 6/154, `intel.alderlake.v1`). Any other CPU, including Cascade Lake and Cooper Lake, fails with `E_PROFILE_HOST_UNKNOWN`. A unit test keeps the copy equal to the profiles of the OpenVMM submodule's pinned revision, which it reads from the gitlink's commit, so the CI jobs validate the NVX CLI after they check out OpenVMM. The host OS's invariant-TSC flags (`constant_tsc nonstop_tsc`, or the CPUID bit on Windows) are recorded as evidence and never fail the check, because they don't decide what a guest observes: Azure WHP hosts show the CPUID bit but cannot offer invariant TSC to partitions, and their guests measure tens of nanoseconds of skew |
| H2 | Vendor, family, model, stepping, microcode, and OS build from `/proc/cpuinfo` or the Windows registry. Then `openvmm --hypervisor <backend> --cpu-fingerprint <path>` writes the host's CPU fingerprint (by default `nvx-cpu-fingerprint-<backend>.json` in the probe directory; `--cpu-fingerprint` overrides it), checks it against the profile that `auto` selects from OpenVMM's catalog, which shares one profile per generation across backends, and prints one `NVX-CPU-PROFILE:` line. H2 reports the generation and the profile from that line, so a profile that an OpenVMM pin adds qualifies its hosts without an NVX change. H2 requires exit status 0, `status=pass`, the same backend, and a catalog profile of the reported generation, reports the profile and surface digests, and otherwise fails with OpenVMM's code, for example `[E_PROFILE_HOST_UNKNOWN]` on a CPU that no profile serves or `[E_PROFILE_UNSUPPORTED]` naming every unsupported CPUID bit. On WHP, OpenVMM checks the entries outside the profile on a probe partition configured from the profile, as a cold boot does. With `--no-openvmm`, H2 maps the host with NVX's copy of the catalog instead: `skylake-sp` (6/85, steppings 0 to 4, `intel.skylake-sp.v1`), `icelake-sp` (6/106, `intel.icelake-sp.v1`), `emeraldrapids` (6/207, `intel.emeraldrapids.v1`), `alderlake` (6/151 and 6/154, `intel.alderlake.v1`), or `milan` (AMD 25/1, `amd.milan.v1`). Any other CPU, including Cascade Lake and Cooper Lake, fails with `E_PROFILE_HOST_UNKNOWN`. A unit test keeps the copy equal to the profiles of the OpenVMM submodule's pinned revision, which it reads from the gitlink's commit, so the CI jobs validate the NVX CLI after they check out OpenVMM. The host OS's invariant-TSC flags (`constant_tsc nonstop_tsc`, or the CPUID bit on Windows) are recorded as evidence and never fail the check, because they don't decide what a guest observes: Azure WHP hosts show the CPUID bit but cannot offer invariant TSC to partitions, and their guests measure tens of nanoseconds of skew |
| H3 | `openvmm --x-time-abi-verify` builds the partition and runs the time ABI preflight without running the guest. Its `NVX-TIME-ABI-VERIFY:` line must report `status=ok` for the backend, plausible declared and native TSC rates, the backend's LAPIC rate, and a `cpu_profile` that is a catalog profile, `<vendor>.<generation>.v<revision>` of any generation but a host profile's `host`, and, when H2 runs too, a revision of the profile H2 names. A failed preflight reports OpenVMM's code, for example `[E_TSC_SYNC_UNSUPPORTED]` |
| H4 | Samples of the TSC against the host's monotonic clocks, with sleeps between them so the host's CPUs idle: 13 samples 10 s apart, or 3 samples 1 s apart with `--ci-schedule`. Each sample reads the TSC between two reads of a clock, keeping the tightest of 64 brackets; its uncertainty is half the bracket plus half the clock's resolution. A clock that returns the same value to consecutive reads is coarser than one read, so the probe takes its smallest step as its resolution: Hyper-V's reference TSC page advances the Linux clocks in 100 ns steps although `clock_getres` reports 1 ns. The interval stability is judged against a clock that time synchronization never steers, `CLOCK_MONOTONIC_RAW` on Linux and `QueryPerformanceCounter` on Windows: every interval between consecutive samples must be conclusive within 0.25 ppm, and the interval rates must agree within 1 ppm. chrony's frequency updates move `CLOCK_MONOTONIC`'s rate by up to several ppm between seconds on the Azure runners, which says nothing about the TSC. The rate over the whole window is measured against the disciplined clock, `CLOCK_MONOTONIC` on Linux, and must lie within 100 ppm of the rate H3 reports when H3 runs in the same invocation. A Linux host's clocksource is recorded as evidence |
| H5 | Pinned-thread ping-pong rounds over every pair of host CPUs; `max_abs_offset_ns` is at most 1,000, the measurement is conclusive, and no pair stalls |
Expand Down
Loading
Loading