Shard 2 of the smoke gate has not completed on any recent run. It is not being cancelled by
anything — it hangs on one suite and hits the 30-minute job timeout, and GitHub badges a
timed-out job as cancelled, which reads as "superseded by a newer run" rather than "this quarter
of the gate did not finish". As a result 13 suites have never executed.
The measurement
Three consecutive runs, shard 2's duration against the other shards:
| run |
sha |
shard 2 |
shard 0 |
shard 1 |
shard 3 |
| 32958247575 |
28f9bf427 |
cancelled after 30m16s |
— |
— |
— |
| 32960622394 |
a67e704b7 |
cancelled after 30m15s |
10m46s |
0m58s |
13m01s |
| 32964962658 |
47b7d025d |
cancelled after 30m |
10m43s |
1m33s |
12m11s |
timeout-minutes: 30 is set at .github/workflows/ci.yml:67. Both completed cases land within
16 seconds of it. fail-fast: false is set at ci.yml:70, so sibling shard failures do not cancel
it, and cancel-in-progress is ${{ github.event_name == 'pull_request' }} (ci.yml:21), so a
push to main does not cancel it either. Nothing cancels shard 2. It runs out of time.
Where it hangs, and that it is deterministic
Both timed-out runs whose logs are retrievable started exactly 83 of 96 suites and stopped on
the same one:
smoke:lang-pins 11:19:37
smoke:runtime-journal-store 11:19:37
smoke:runtime-mesh-wait 11:19:39
smoke:runtime-fork 11:20:03
smoke:agui-opencode-source 11:20:04
smoke:opencode-events-release 11:20:05 <- still running at 11:27:03 when the job was killed
Its five predecessors each took between under a second and 24 seconds. smoke:opencode-events-release
(extensions/connector-opencode/smoke/events-release.smoke.ts) had been running 7 minutes and
had not finished. The same suite, at the same position, on run 32958247575 as well. This is a
deterministic hang, not a slow shard: shard 2 reached suite 83 in 23 minutes and then one suite
consumed the rest.
On the kill, the runner terminates orphans including two nats-server processes, so the suite is
leaving a live broker behind when it wedges.
What has never run
The 13 suites after the hang point on shard 2, none of which has executed in any observed run:
smoke:agui-codex-map smoke:hermes-launch-env smoke:overflow-churn-bound
smoke:web-snapshot smoke:history-flood smoke:empty-id-ingest
smoke:core-history-limit smoke:remote-exchange:live smoke:control-transport-dial
smoke:lang-engine smoke:jcode-provider-disconnect smoke:no-implicit-general
smoke:int2-revoke:live
Arithmetic cross-checks against CI's own numbers: the file parses to 384 suites, shard 2 holds 96,
and the hang suite is #83 — matching the job's shard 2/4 — 96 of 384 smokes header and its 83
started.
Two of those are worth calling out because they guard behaviour that shipped recently and were
presumably added to prevent exactly the regressions they cannot currently catch:
smoke:no-implicit-general (the implicit general channel floor removal) and
smoke:control-transport-dial (the transport-aware control dial).
Why this stayed invisible
A hang and a cancellation produce the same badge. Anyone reading the checks list sees three shards
with real verdicts and one greyed-out cancelled, which is what a superseded run looks like. There
is no signal anywhere in the UI that a quarter of the gate timed out, and the log that would show it
is only published once the job ends — so an in-flight investigation cannot see it either.
Suggested scope
- Fix or quarantine the hang in
smoke:opencode-events-release. It should fail loudly on its own
budget rather than consuming the shard's.
- Give suites an individual timeout so one wedged suite cannot eat the shard, and make the shard
runner report which suite it was in when killed.
- Treat a
cancelled smoke shard as a failure rather than as noise — as it stands, the gate cannot
distinguish "did not run" from "was superseded", which is the same class as a check that cannot
fail.
Item 2 is the one that prevents recurrence; item 1 only fixes today's instance.
Shard 2 of the smoke gate has not completed on any recent run. It is not being cancelled by
anything — it hangs on one suite and hits the 30-minute job timeout, and GitHub badges a
timed-out job as
cancelled, which reads as "superseded by a newer run" rather than "this quarterof the gate did not finish". As a result 13 suites have never executed.
The measurement
Three consecutive runs, shard 2's duration against the other shards:
28f9bf427a67e704b747b7d025dtimeout-minutes: 30is set at.github/workflows/ci.yml:67. Both completed cases land within16 seconds of it.
fail-fast: falseis set atci.yml:70, so sibling shard failures do not cancelit, and
cancel-in-progressis${{ github.event_name == 'pull_request' }}(ci.yml:21), so apush to main does not cancel it either. Nothing cancels shard 2. It runs out of time.
Where it hangs, and that it is deterministic
Both timed-out runs whose logs are retrievable started exactly 83 of 96 suites and stopped on
the same one:
Its five predecessors each took between under a second and 24 seconds.
smoke:opencode-events-release(
extensions/connector-opencode/smoke/events-release.smoke.ts) had been running 7 minutes andhad not finished. The same suite, at the same position, on run 32958247575 as well. This is a
deterministic hang, not a slow shard: shard 2 reached suite 83 in 23 minutes and then one suite
consumed the rest.
On the kill, the runner terminates orphans including two
nats-serverprocesses, so the suite isleaving a live broker behind when it wedges.
What has never run
The 13 suites after the hang point on shard 2, none of which has executed in any observed run:
Arithmetic cross-checks against CI's own numbers: the file parses to 384 suites, shard 2 holds 96,
and the hang suite is #83 — matching the job's
shard 2/4 — 96 of 384 smokesheader and its 83started.
Two of those are worth calling out because they guard behaviour that shipped recently and were
presumably added to prevent exactly the regressions they cannot currently catch:
smoke:no-implicit-general(the implicitgeneralchannel floor removal) andsmoke:control-transport-dial(the transport-aware control dial).Why this stayed invisible
A hang and a cancellation produce the same badge. Anyone reading the checks list sees three shards
with real verdicts and one greyed-out
cancelled, which is what a superseded run looks like. Thereis no signal anywhere in the UI that a quarter of the gate timed out, and the log that would show it
is only published once the job ends — so an in-flight investigation cannot see it either.
Suggested scope
smoke:opencode-events-release. It should fail loudly on its ownbudget rather than consuming the shard's.
runner report which suite it was in when killed.
cancelledsmoke shard as a failure rather than as noise — as it stands, the gate cannotdistinguish "did not run" from "was superseded", which is the same class as a check that cannot
fail.
Item 2 is the one that prevents recurrence; item 1 only fixes today's instance.