You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
ci: the azure-mshv-scus runners fail the MSHV performance gate on guest-exit teardown and 512 MiB restore #415
The performance gate fails whenever the Platform / Linux / MSHV / Virtual machine job runs on azure-mshv-scus-1 or azure-mshv-scus-z1-1, whatever the pull request changes. Those runners measure MSHV guest-exit teardown about 15 to 29 ms slower, and the 512 MiB shell snapshot restore at 4 and 8 vCPUs about 20 to 40 ms slower, than the runners that measured the baseline window (azure-mshv-1, -5, -6, and -7). Every rerun that lands on another MSHV runner passes the gate.
#394 and #404 reported the same pattern: those runners use another scale set and host kernel (6.6.137.mshv2-2.azl3, against 6.6.148-200.azl3). Both runners also run Rust 1.99.0, so OpenVMM vmm-tests / Linux / MSHV fails there too (#402).
Evidence
The MSHV platform job's runner and the gate's result in recent pull-request runs:
The same metrics fail in each of the five failing attempts of #411 and #412:
openvmm_cold_start_guest_exit_teardown: 40 to 48 ms, against a base median of 24.94 ms.
openvmm_snapshot_restore_guest_exit_teardown: 30 to 37 ms, against 8.16 ms.
shell_snapshot_restore_512_mib: 69 to 71 ms at 4 vCPUs (base 46.25 ms), and 91 to 95 ms at 8 vCPUs (base 52.95 ms).
Effect
Any MSHV runtime pull request fails the gate when its platform job lands on these runners. CI routes by labels, which all MSHV runners share, so the only remedy is to rerun until the job lands elsewhere. The gate cannot tell a regression from runner placement. dev pushes that ran there, such as 37503416598 for #404's merge, may also feed the baseline history with these slower samples.
Options
Give the scus runners a series of their own, as doc/benchmarks.md describes for AMD runners, or keep the platform jobs off them, for example with an extra runner label for the gate's platform jobs.
Or align their host kernel and scale set with the baseline runners, and confirm that their measurements then match.
Note
Runner naming update (2026-10-07): Names in linked historical jobs and dated tables remain exactly as GitHub reported them. Current sequential aliases and runner registrations are mapped by VM instance in #211's rename note. Use the current name when connecting to or scheduling a runner.
Summary
The performance gate fails whenever the
Platform / Linux / MSHV / Virtual machinejob runs onazure-mshv-scus-1orazure-mshv-scus-z1-1, whatever the pull request changes. Those runners measure MSHV guest-exit teardown about 15 to 29 ms slower, and the 512 MiB shell snapshot restore at 4 and 8 vCPUs about 20 to 40 ms slower, than the runners that measured the baseline window (azure-mshv-1,-5,-6, and-7). Every rerun that lands on another MSHV runner passes the gate.#394 and #404 reported the same pattern: those runners use another scale set and host kernel (
6.6.137.mshv2-2.azl3, against6.6.148-200.azl3). Both runners also run Rust 1.99.0, soOpenVMM vmm-tests / Linux / MSHVfails there too (#402).Evidence
The MSHV platform job's runner and the gate's result in recent pull-request runs:
azure-mshv-scus-1azure-mshv-7azure-mshv-1azure-mshv-5azure-mshv-1azure-mshv-scus-z1-1, thenazure-mshv-scus-1twiceazure-mshv-6azure-mshv-scus-1The same metrics fail in each of the five failing attempts of #411 and #412:
openvmm_cold_start_guest_exit_teardown: 40 to 48 ms, against a base median of 24.94 ms.openvmm_snapshot_restore_guest_exit_teardown: 30 to 37 ms, against 8.16 ms.shell_snapshot_restore_512_mib: 69 to 71 ms at 4 vCPUs (base 46.25 ms), and 91 to 95 ms at 8 vCPUs (base 52.95 ms).Effect
Any MSHV runtime pull request fails the gate when its platform job lands on these runners. CI routes by labels, which all MSHV runners share, so the only remedy is to rerun until the job lands elsewhere. The gate cannot tell a regression from runner placement.
devpushes that ran there, such as 37503416598 for #404's merge, may also feed the baseline history with these slower samples.Options
scusrunners a series of their own, asdoc/benchmarks.mddescribes for AMD runners, or keep the platform jobs off them, for example with an extra runner label for the gate's platform jobs.Note
Runner naming update (2026-10-07): Names in linked historical jobs and dated tables remain exactly as GitHub reported them. Current sequential aliases and runner registrations are mapped by VM instance in #211's rename note. Use the current name when connecting to or scheduling a runner.