Skip to content

ci: the azure-mshv-scus runners fail the MSHV performance gate on guest-exit teardown and 512 MiB restore #415

Description

Summary

The performance gate fails whenever the Platform / Linux / MSHV / Virtual machine job runs on azure-mshv-scus-1 or azure-mshv-scus-z1-1, whatever the pull request changes. Those runners measure MSHV guest-exit teardown about 15 to 29 ms slower, and the 512 MiB shell snapshot restore at 4 and 8 vCPUs about 20 to 40 ms slower, than the runners that measured the baseline window (azure-mshv-1, -5, -6, and -7). Every rerun that lands on another MSHV runner passes the gate.

#394 and #404 reported the same pattern: those runners use another scale set and host kernel (6.6.137.mshv2-2.azl3, against 6.6.148-200.azl3). Both runners also run Rust 1.99.0, so OpenVMM vmm-tests / Linux / MSHV fails there too (#402).

Evidence

The MSHV platform job's runner and the gate's result in recent pull-request runs:

Run Pull request MSHV platform runner Gate
37410478636 #401 azure-mshv-scus-1 failure
37419824754 #401 azure-mshv-7 success
37409832176 #394 azure-mshv-1 success
37439708257 #404 azure-mshv-5 success
37448382849 #404 azure-mshv-1 success
37518293095, attempts 1 to 3 #411 azure-mshv-scus-z1-1, then azure-mshv-scus-1 twice failure
37518293095, attempt 4 #411 azure-mshv-6 success
37520322807, attempts 1 and 2 #412 azure-mshv-scus-1 failure

The same metrics fail in each of the five failing attempts of #411 and #412:

  • openvmm_cold_start_guest_exit_teardown: 40 to 48 ms, against a base median of 24.94 ms.
  • openvmm_snapshot_restore_guest_exit_teardown: 30 to 37 ms, against 8.16 ms.
  • shell_snapshot_restore_512_mib: 69 to 71 ms at 4 vCPUs (base 46.25 ms), and 91 to 95 ms at 8 vCPUs (base 52.95 ms).

Effect

Any MSHV runtime pull request fails the gate when its platform job lands on these runners. CI routes by labels, which all MSHV runners share, so the only remedy is to rerun until the job lands elsewhere. The gate cannot tell a regression from runner placement. dev pushes that ran there, such as 37503416598 for #404's merge, may also feed the baseline history with these slower samples.

Options

  • Give the scus runners a series of their own, as doc/benchmarks.md describes for AMD runners, or keep the platform jobs off them, for example with an extra runner label for the gate's platform jobs.
  • Or align their host kernel and scale set with the baseline runners, and confirm that their measurements then match.

Note

Runner naming update (2026-10-07): Names in linked historical jobs and dated tables remain exactly as GitHub reported them. Current sequential aliases and runner registrations are mapped by VM instance in #211's rename note. Use the current name when connecting to or scheduling a runner.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions