Skip to content

Recalibrate Windows warm P50 budget for schema-v2 variance #513

Description

Problem

PR #512 changes only warm P95 budgets, but its first Windows performance job failed the unchanged warm full-refresh P50 gate:

  • exact schema-v2 baseline ad7ca14: 105ms
  • schema-v2 PR measurements for unchanged product code: 117ms, 133ms, 160ms, 164ms, and 182ms
  • existing budget: 50ms / 30%, which blocks above 155ms for that baseline

Three observed PR runs therefore exceed the effective ceiling produced by the unusually low merged baseline, despite stable warm P95, cold P50, inventory, and timeout diagnostics. The isolated #512 rerun passed at 117ms, confirming runner variance rather than a product change.

Scope

After #512 merges:

  • calibrate the Windows warm full-refresh P50 dual budget from repeated schema-v2 hosted runs;
  • keep Linux/macOS P50, all P95, cold-refresh, startup, coverage, and schema behavior unchanged unless independent evidence requires otherwise;
  • add tests proving observed Windows variance passes and a material median regression fails;
  • update docs/QUALITY_SNAPSHOTS.md with the calibration source.

Acceptance criteria

  • repeated unchanged-head Windows schema-v2 runs stay green;
  • a sustained Windows warm median regression fails a tested gate;
  • Linux/macOS and cold/tail gates are unchanged;
  • full CI, coverage, performance, CodeQL, and comparator tests pass.

Depends on #512.

Activity

  1. karthiknadig commented on Aug 11, 2026

    @karthiknadig
    MemberAuthor

    Windows schema-v2 warm full-refresh P50 evidence:

    Observed range: 105-182ms (77ms). The existing 50ms/30% budget blocks above 155ms against the exact baseline, so three unchanged-product runs fail.

    Proposed Windows-only warm P50 budget: 150ms / 50%. Against the 105ms baseline it blocks above 255ms, leaving nearly 2x the observed absolute range as headroom while still rejecting a sustained ~2.5x median regression. Linux/macOS, all P95, cold, startup, coverage, and schema behavior remain unchanged.

  2. added a commit that references this issue on Aug 11, 2026
  3. karthiknadig commented on Aug 11, 2026

    @karthiknadig
    MemberAuthor

    Final quality conclusion for PR #515 at 77eb4ce:

    • Exact-head Copilot review: 3/3 files, 0 comments, 0 unresolved threads.
    • Full CI, lint, CodeQL, Linux/Windows coverage, and all exact-base performance gates pass.
    • Windows warm P50/P95: 163/180ms vs 137/144ms baseline; cold P50 157ms vs 138ms.
    • Linux warm P50/P95: 59/62ms vs 60/62ms; macOS 93/107ms vs 112/137ms.
    • All artifacts contain 10 samples, matching inventories, and empty interpreter-timeout maps.
    • Line/function coverage is unchanged on Linux and Windows.
    • Comparator suite: 35/35 passed; observed 182ms Windows variance passes, 300ms fails, and explicit dual-budget coverage remains.
    • The initial Linux check failed closed because the freshly merged exact-base baseline artifact was not yet available; its benchmark was healthy and the rerun passed after baseline completion.
    • The earlier manual calibration run also passed after retrying a transient self-signed-certificate failure in the baseline-download action.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestimportantIssue identified as high-priority

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions