Skip to content

Windows API lane hangs in the MCP concurrency tests and has cancelled every nightly since 20 August #1762

Description

@sunt05

Every scheduled run of build-publish_to_pypi.yml from 20 August to 7 September (19 runs, e.g. 34075203398, 34005586564, 33938085189) ends cancelled after about an hour. In each run the three Windows API cross-CPython tests jobs (cp312, cp313, cp314) reach the 45-minute timeout-minutes in test-api-cross-python-reusable.yml (#1563) while every Linux and macOS job passes.

The Windows cp312 log of run 34075203398 reports 1849 test results, the last being test/mcp/test_protocol_handshake.py::test_prompts_advertise_all_three at 02:49:58, and then nothing until the cancellation at 03:14:17. The next node in collection order is test_concurrent_query_knowledge_does_not_block_event_loop, followed by test_concurrent_query_knowledge_and_search_schema. Both are marked slow, so they run only in the all tier, and both pass on Linux and macOS in the same run in 2.5 to 4.7 s. #1709 (19 August) is what made them run against the installed suews-mcp executable; the last green nightly is 19 August. The cp314 Windows job stops at the same test.

Consequences so far:

  • No TestPyPI dev build since 2026.8.19.dev0.
  • No Windows API coverage for 19 days: the all tier on Windows (cp312, cp313, cp314) is also the only place the UMEP/QGIS lane runs. The Linux and macOS API jobs completed green in the same runs.
  • Physics coverage is unaffected: -m "physics and slow and not core" collects nothing, so the nightly's physics tier selects the same 166 tests as the merge queue.
  • Nothing reports a failed scheduled run; the observability job runs only on pull_request and merge_group.

This is a defect on Windows, not a case for a platform skip or a longer cap: the same two tests finish in seconds on the other platforms, and Windows API coverage is coverage the project wants. Two things to fix, and one to add:

  1. The hang. Reproduce on a Windows workflow_dispatch with tier all and only the Windows platform. Both tests spawn the suews-mcp server over stdio and gather two tool calls, so the candidates are the stdio client on the Windows proactor event loop, a server subprocess that does not exit so the session teardown blocks, or the worker thread the tool bodies offload to (fix(mcp): offload sync tool bodies to a worker thread (closes #1412) #1439) never returning on Windows. Fix the cause in suews_mcp or the test harness.
  2. The wider Windows time gap. On identical node sets Windows takes 789 s where Linux takes 541 s in the merge queue, and 1206 to 1324 s where Linux takes 860 to 1075 s in the last green nightly. Profile the Windows lane (subprocess spawn cost, temporary-directory I/O, path resolution) and fix what is found rather than budgeting for it; the 45-minute cap should stop being load-bearing.
  3. Detection. A per-test timeout via pytest-timeout (thread method; the signal method is Unix-only) so a hang is reported with the test's name instead of hidden by a job cancellation, and a final job on schedule runs that opens or updates a pinned issue when any job fails or is cancelled and closes it on the next green run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    1-bugSomething isn't working2-infra:ciCI/CD pipelines, GitHub Actions2-infra:testTesting infrastructure, pytest

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions