Skip to content

ci: qualify runners without an invariant TSC by measurement - #432

Merged
Pedro Henrique Penna (ppenna) merged 2 commits into
devfrom
agents/azure-mshv-2-cpu-fingerprint-fixes
Oct 8, 2026
Merged

Pedro Henrique Penna (ppenna) merged 2 commits into
devfrom
agents/azure-mshv-2-cpu-fingerprint-fixes

Conversation

@ppenna

Copy link
Copy Markdown
Contributor

CI no longer fails a runner because its host OS lacks nonstop_tsc. Runners without an invariant TSC now qualify on what their guests measure, and the platform jobs run the guest warp probe before their benchmarks. This lets azure-mshv-2 back into rotation: its mshv label can return once this merges.

Refs #263, #265.

Why

azure-mshv-2 (esaurez-nvx-ci-linux-vmss000001, Xeon Platinum 8370C, Ice Lake-SP) lost its mshv label in #263. Its host OS has no nonstop_tsc, and before the time ABI its guests hit cross-vCPU TSC warps (#211, #265). The time ABI (#325) replaced those checks with the guest warp probe and its 1 µs bound, and the doctor already treats the invariant-TSC flags as evidence only.

But validate-runner and setup-linux-runner.sh still failed every host without nonstop_tsc. The spec kept that gate "until every job that runs guests also runs H6". Re-adding the label alone would therefore fail every job on the runner at Validate host TSC.

The CPU fingerprint fixes (#411, #412) don't change this host's eligibility. The pre-#411 OpenVMM (07850d27a) and #412's pin (3951c8415) write the same fingerprint there, and both pass intel.icelake-sp.v1 with the same surface digest.

Commits

Commit Change
ee2295f9 ci: qualify runners without an invariant TSC by measurement validate-runner reports the TSC flags and only notes a missing nonstop_tsc. A new warp-probe input adds H6 on CI's schedule, and the platform jobs set it, which adds about 25 s before their benchmarks. The microVM scenarios already probe after every boot and restore. The OpenVMM vmm-tests boot OpenVMM's own guests, whose verdicts don't depend on the skew bound. setup-linux-runner.sh warns instead of stopping. The tests, the CI guide, the time ABI spec, the setup README, and Copilot setup's notice follow
b55765f3 doctor: give H4's probe its sampling window on top of --timeout The long H4 schedule sleeps 120 s across its 13 samples, which the default 120 s --timeout didn't cover, so nvx.py doctor --backend <backend> timed out H4

Validation

On azure-mshv-2, with release v0.1.0-dev.c0f099d0ca91 (#412's merge) and its source:

Check Result
nvx.py doctor H1 to H7 on the long schedules Pass. H6 measured at most 26 ns, H5 38 ns, and H4's interval rates agree within 0.006 ppm
--cpu-fingerprint with OpenVMM before and after #411 Both pass intel.icelake-sp.v1, with the same surface digest
CI's MSHV microVM job, replayed step by step: the default suite, managed exec config, Ubuntu, Azure Linux, the 5 sandbox and share steps, and aci_edge_sandboxes Pass
CI's MSHV debug-kernel microVM job Pass
10 more rounds of smp, smp-snapshot, and restore-processors Pass. With the two jobs above, 288 warp probe runs measured at most 36 ns of offset and 4 ns of backward step
3 CI-equivalent benchmark runs, replayed through the performance gate against dev's history Pass: 37 metrics, no regression. Cold starts are about 10% faster than the fleet median. Guest-exit teardown and the 512 MiB restores are at most 1.5% slower, so there's no #415 pattern. The worst deltas are UDP round-trip ops (8 to 12% lower) and snapshot generation (+0.35 ms)
doctor --checks H1 H2 H4 H6 H3 --ci-schedule, the platform jobs' new checks Pass, 3 of 3. H6 adds about 24 s
The new Report host TSC step and setup-linux-runner.sh --check-only Pass, with the notice and the warning
doctor H1 to H7 at the default --timeout, with the fix Pass in 183 s; H4 timed out before

Before the time ABI, every vmm-tests, platform, and unit test job that ran on azure-mshv-2 (September 23 to 29) passed its test steps: 13, 16, and 13 runs. Only the NVX microVM restore scenarios failed there.

Repository checks on this head:

  • On Linux, the validate-nvx unit tests (753), adversarial tests (70), host-connect tests (4), and every --help check pass.
  • On Windows, the unit tests pass except test_materialize_kernel_provenance_inputs_bypasses_mutable_index. It fails identically on dev there because the local Git signing configuration can't sign the fixture's commit.
  • Ruff check and format, Pyright for Linux and Windows, compileall, and the pinned shellcheck and shfmt pass.

Left to CI: the platform jobs' new H6 on the KVM, MSHV, and WHP runners, and the rest of the matrix.

After merge

  • Re-add the runner's label: gh api -X POST repos/microsoft/nvx/actions/runners/51/labels -f 'labels[]=mshv'.
  • My account is now in azure-mshv-2's mshv group, as on azure-mshv-5, so developers can run guests there.

azure-mshv-2, an Azure Xeon Platinum 8370C (Ice Lake-SP) MSHV runner
whose host OS lacks nonstop_tsc, has been out of rotation since #263:
before the time ABI, its guests hit cross-vCPU TSC warps (#211, #265).
The time ABI replaced those warp-sensitive checks with the guest warp
probe and its 1 us bound, and the doctor already treats the
invariant-TSC flags as evidence only. Yet validate-runner and the Linux
runner setup script still failed every host without nonstop_tsc, a gate
meant to stay until every job that boots guests also runs H6.

- validate-runner reports the host's TSC flags and only notes a missing
  nonstop_tsc. A new warp-probe input adds H6, on CI's schedule, to the
  doctor for jobs that boot guests without the microVM scenarios' own
  probes.
- The platform jobs set it, which adds about 25 s before their
  benchmarks. The microVM scenarios already probe after every boot and
  restore.
- setup-linux-runner.sh warns instead of stopping and points to the
  doctor.
- The OpenVMM vmm-tests stay without a probe: they boot OpenVMM's own
  test guests, whose verdicts don't depend on the skew bound. Before the
  time ABI, every vmm-tests, platform, and unit test job that ran on
  azure-mshv-2 passed its test steps (13, 16, and 13 runs); only the
  NVX microVM restore scenarios failed there.
- Copilot setup's notice no longer predicts unstable restores.

On azure-mshv-2, with release v0.1.0-dev.c0f099d0ca91:

- nvx.py doctor passes H1 to H7 on the long schedules, and H6 measured
  at most 26 ns. The CPU profile check passes with the same surface
  digest on OpenVMM before and after #411's MSHV fix.
- The CI MSHV microVM jobs pass on the production and debug kernels, as
  do ten more rounds of smp, smp-snapshot, and restore-processors: 288
  warp probe runs measured at most 36 ns of offset and 4 ns of backward
  step.
- Three CI-equivalent benchmark runs pass the performance gate against
  dev's history, with no regression in 37 metrics.
- The Report host TSC step, setup-linux-runner.sh --check-only, and the
  platform jobs' doctor checks with H6 pass there.

The runner's mshv label can return once this merges.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
On the qualification schedule, H4's probe sleeps 10 s between each of
its 13 samples, 120 s in all, but run_probe allowed it only --timeout,
whose default is also 120 s. `nvx.py doctor --backend <backend>`, the
documented way to qualify a host, then failed H4 with a timeout whenever
taking the samples outlasted that margin, as it did on azure-mshv-2.

run_probe now takes the time the probe spends sleeping by design and
allows it on top of --timeout, and H4 passes its sampling window. CI's
short schedule gains only its 2 s of sleeps. The --timeout help and the
usage reference say so.

On azure-mshv-2, the full qualification at the default --timeout now
passes H1 to H7 in 183 s.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 8, 2026 02:34

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Relaxing runner eligibility needs maintainer sign-off after the new H6 checks pass across KVM, MSHV, and WHP CI.

0 open findings

What changed in this PR

Replaces the invariant-TSC runner requirement with measured guest qualification, supporting the return of azure-mshv-2 to CI.

Changes:

  • Reports missing invariant-TSC support instead of rejecting runners.
  • Adds H6 guest warp checks before platform benchmarks.
  • Extends H4’s timeout by its sampling window, with regression tests and documentation updates.
File Description
scripts/​test_time_abi.py Tests sampling-window timeout allowances.
scripts/​test_nvx_tools.py Tests relaxed eligibility and H6 wiring.
scripts/​setup/​setup-linux-runner.sh Warns instead of rejecting missing invariant TSC.
scripts/​setup/​README.md Documents measurement-based runner qualification.
scripts/​nvx_tools/​doctor.py Adds H4’s sampling window to its timeout.
doc/​usage.md Clarifies timeout behavior.
doc/​design/​time-abi.md Updates qualification policy and measured evidence.
doc/​ci.md Documents revised CI qualification checks.
.github/​workflows/​run-platform.yml Enables H6 before benchmarks.
.github/​workflows/​copilot-setup-steps.yml Updates the missing-TSC notice.
.github/​actions/​validate-runner/​action.yml Removes TSC rejection and adds optional H6 checks.

🧠 Review effort: Balanced


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

@ppenna
Pedro Henrique Penna (ppenna) merged commit f813dc7 into dev Oct 8, 2026
31 checks passed
@ppenna
Pedro Henrique Penna (ppenna) deleted the agents/azure-mshv-2-cpu-fingerprint-fixes branch October 8, 2026 03:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants