Repository navigation
ci(ios): rerun a module once when the browser never reached the suite; boot after setup - #495
Conversation
… suite The 90s iOS flow timeout from #469 was not enough on main (run 37395358046, Config RP, oidcc-client-test-discovery-jwks-uri-keys: no user after 94608ms, NO authorization request received). The simulator log shows why no timeout can fix that case: the SafariViewService process hosting the session went unresponsive at 01:12:40 and was killed by the watchdog at 01:13:13 (exit 0x8badf00d). The session never recovered, and the cancel came at +94.65s. In other runs the stall was a WebKit WebContent launch taking 27-56s. Pre-warming once is not a fix either. Every session cold-launches its own WebContent process (93 launches for 94 sessions in that run), so there is no long-lived process to keep warm, and the stalls land mid-plan, not only on the first session. So when, and only when, the suite's own log proves it received NO authorization request for an instance, the harness stops that instance and runs the module once more on a fresh one. The suite observed nothing there and judged nothing, so no verdict can be hidden. It never reruns when any authorization request arrived (every negative module), never reruns a rerun, and never reruns on an unreadable log. The rerun is printed to the job log and carried into the module's verdict line. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
The open question for the iOS browser stalls (#469) is whether a dialog blocks the login. The existing capture could not answer it: it only streamed Runner and SafariViewService, and nothing recorded what was on screen. * While a login is still pending 20s, 50s and 80s in (a healthy iOS login takes ~3-5s), the harness reads the native UI tree of SpringBoard and of the app through Patrol's iOS automator. It prints any alert or sheet with its texts as "[e2e] STALL-PROBE". A once-per-plan "[e2e] SCREEN-PROBE-SELFTEST" proves the probe can see native UI, so a "no alert" during a stall means something. * The iOS job streams a second, info-level log of SpringBoard, backboardd, runningboardd, and any alert-related lines. It takes a simulator screenshot at every STALL-PROBE marker and a final one at the end. Logs are gzipped and uploaded with the screenshots in ios-simulator-log. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
… synchronous Review of run 37423858857 found the first probe could not rule a dialog out: * With a null selector, Patrol's iOS automator snapshots the FOREGROUND app and ignores appId (IOSAutomator.getUITreeRoots). So "springboard" was really the app again, and every probe reported only application:'Example'. The probe now runs selector queries, which honour the bundle: SpringBoard alerts, app alerts, and the browser sheet (its toolbar shows "certification.openid.net"). Each query gets its own deadline and timing, so a hung query is named. * Positive controls. SHEET-PROBE-CONTROL probes 1.5s into each plan's first healthy login, while the browser sheet is up. ALERT-PROBE-CONTROL (Basic RP, once) opens a non-ephemeral ASWebAuthenticationSession to example.com, whose "Wants to Use … to Sign In" consent alert is a known SpringBoard alert. The probe must report it before flowTimeoutSeconds cancels the session. * Screenshots were up to a minute late because the watcher grepped the buffered log stream. The harness now drops a marker file in the app's tmp/, which a host watcher polls every 0.5s. Each screenshot is named after the harness's UTC timestamp and tag, and has a .txt recording when the marker was seen and when the screenshot started and ended. * host-load.log samples the macOS host every 2s (load, top CPU, memory, disk). Inside the simulator, process launches were seen stalling for 30-150s with launches queueing up behind them; this shows whether the host was starved at that moment. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
Run 37428137698 lost Basic, Config, Hybrid and Implicit RP to the probe
itself. The app-side queries used Patrol's default appId, which comes
from pubspec's patrol_plus ios bundle_id (com.bdayadev.oidc.example).
The iOS Runner target actually builds com.bdayadev. XCTest recorded
"Failed to resolve query: Application com.bdayadev.oidc.example is not
running" as a test failure, tore the test down and SIGTERMed the app
mid-plan (runningboardd: termination reported by launchd (2, 15, 15)).
* The probe now names com.bdayadev explicitly.
* The positive controls move out of the plans into their own patrolTest
("zz iOS native probe positive controls", sorted after every plan).
They poll once a second for up to 20s instead of taking one snapshot,
and report FOUND / NOT FOUND with a timed poll history:
- SHEET-PROBE-CONTROL: an ephemeral session on
certification.openid.net; "browser sheet" must match.
- ALERT-PROBE-CONTROL: a non-ephemeral session on example.com, whose
sign-in consent alert one of the alert queries must report.
* The iOS job records the simulator screen continuously (simctl io
recordVideo). The marker-triggered screenshots were 13s late to start
and took 37s in that run. The host sampler uses ps instead of top,
every 3s, so it does not compete for the 3 vCPUs it is measuring.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
…cording
The probe could not be made safe. In run 37430662455, with the bundle id
fixed, it still failed Hybrid RP and its own control test. That is 64
XCTest failures in a run where every Dart test passed. While a browser
sheet is up, XCUITest queries on the app take 8s+ or time out. A query
that matches nothing ("Failed to get matching snapshot: No matches
found") is recorded as a test failure. A diagnostic that changes the
outcome it observes is worse than none.
What answers the dialog question instead is the simulator screen
recording that run made. During the oidcc-client-test-missing-iat stall
(login 08:01:04.1 -> 08:01:48.8, 45s), frames at +2, +10, +20 and +30s
all show the ASWebAuthenticationSession sheet on
"certification.openid.net" with a blank page and no dialog. The sheet is
gone by +40s.
* Patrol test file back to main's. The harness only prints
"[e2e] LOGIN-PENDING <module> after 20/50/80s" as timestamps for
reading the recording.
* The marker screenshot watcher is gone. The screen recording stays.
* The app's log stream drops to info level. At debug, "log" plus
diagnosticd took 40-60% CPU on a host already at load 200-500.
* host-load.log adds the top RSS processes, to name what drives the
swapping (swapouts 0 -> 152744 in that run).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
…ost load The iOS browser stalls (#469) are slow process launches on a starved host. In run 37430662455 the 1-minute load average was 200-500, free memory about 4k pages, and swapouts climbed 0 -> 152744. The screen recording ruled out a dialog. Diagnostics that add load are now opt-in, and load reduction is a switch that can be measured on and off. * Default diagnostics cost nothing during the tests beyond a 3s ps/vm_stat sampler. The browser-flow evidence is read afterwards from the simulator's persisted unified log ("log show"): the harness's flutter lines and runningboardd's WebContent launch lines. .github/scripts/ios_ci_metrics.py turns them into login-duration and launch-latency distributions plus host load/free/swap figures, printed as [ios-metrics] lines. * workflow_dispatch input ios_heavy_diagnostics (default off) brings back the live log streams and the screen recording. * workflow_dispatch input ios_reduce_load (default on; always on for push/PR): - turns host Spotlight indexing off with mdutil (mdworker_shared and spotlightknowledged.updater were top host CPU users); - shuts down any other booted simulator first; - prefers the smallest iPhone the runtime offers over a 1206x2622 Pro-class device, to lighten SimRenderServer/SimMetalHost. GitHub's standard macos-26 arm64 runner is 3 vCPU / 7 GB RAM. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
ReportCrash held 210 MB RSS and CPU mid-run in run 37435756354, on a host that was already swapping. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
Under bash -e, `vpid="$(cat missing-file)"` takes cat's exit status and ended the step (run 37440456274: every test passed, the iOS job failed, and no diagnostics were uploaded). Dry-run of the step under bash -e with xcrun stubbed now exits 0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
Rebuilt on main's 2-shard iOS job. The harness (packages/oidc/example) is taken from main as-is, so the earlier probes, controls and the module rerun leave this branch's diff against main. What stays: * .github/scripts/ios_ci_metrics.py, plus cheap default diagnostics: a 3s ps/vm_stat host sampler while the tests run. The app's lines and runningboardd's WebContent launch lines are read back from the simulator's persisted log afterwards. This replaces main's debug-level live log stream, which itself cost 40-60% CPU on a starved host. Heavy diagnostics (live streams, screen recording) are opt-in via workflow_dispatch. * IOS_REDUCE_LOAD (default on): Spotlight indexing off, other simulators shut down, smallest iPhone preferred. * ios_runner dispatch input, to measure macos-26 (3 vCPU / 7 GB arm64) against macos-26-intel (4 vCPU / 14 GB x64, same Xcode 26.6 and iOS 26.4 runtime). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
Booting alongside the Flutter install slowed that install 3-4x on the 3-vCPU arm64 runner. The Flutter setup script took 65s with a synchronous boot and 178-273s with a concurrent one; cache-pub's hashFiles took 0.5s vs 27-125s. That is a net loss of 80-240s per iOS leg after #492. The boot now overlaps only flutter precache and the Patrol CLI activation, which moves ahead of the wait. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
… suite Re-applied onto main's harness (after #492), unchanged in behaviour from 3903937. Load reduction did not make the iOS simulator's browser stalls go away: they stayed intermittent across every runner configuration measured. When, and only when, the suite's own log proves it received NO authorization request for an instance, the harness stops that instance and runs the module once more on a fresh one. The suite observed nothing there and judged nothing, so no verdict can be hidden. It never reruns when any authorization request arrived (every negative module), never reruns a rerun, and never reruns on an unreadable log. The rerun is printed to the job log and carried into the module's verdict line. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0126i8ymDVXbifp6LrbCiyQL
…help Measured over 2+ iOS runs each, none of these made the simulator browser stalls (#469) go away: * Spotlight off, other simulators shut down, smallest iPhone. Logins >30s per run: 1 and 4 without, 0 and 1 with (single job). In the sharded layout: 0/0 without, 2/1 with. * macos-26-intel (4 vCPU / 14 GB): Set up Environment took 9.5-15 min, and both shards hit the 60-minute timeout in their tests. Both are removed. What stays from the experiment is the cheap default diagnostics, heavy diagnostics as a dispatch input, and the boot ordering. The flaky stall itself is handled by the harness's evidence-gated single rerun. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016g5xVp7yzwhi7f9tLwcNBj
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (5)
Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe conformance runner now retries eligible modules with a fresh suite instance. The iOS workflow adds manual retry controls and collects simulator diagnostics, including login, WebContent launch, and host-load metrics. ChangesiOS conformance runs
Priority: ⬇️ Low Estimated code review effort: 4 (Complex) | ~45 minutes Change: Bug fix Sequence Diagram(s)sequenceDiagram
participant Runner as Conformance runner
participant Suite as Suite instance
participant Predicate as shouldRerunModuleOnFreshInstance
participant Retry as Fresh suite instance
Runner->>Suite: Run module attempt
Suite-->>Runner: Return no user, WAITING status, and suite log
Runner->>Predicate: Check login state, attempt, status, and log
Predicate-->>Runner: Return rerun eligibility
Runner->>Suite: Cancel and clean up discarded instance
Runner->>Retry: Create instance for retry
Merge Risk: ⚪ Minimal · up to This change makes iOS conformance CI retry a module once when the simulator browser never reached the test suite. It also moves simulator boot later in the job and adds diagnostics. Production code is not affected, and no merge-blocking issue was found. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 6 functions across 1 files. (4 skipped: 4 unsupported.)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #495 +/- ##
==========================================
- Coverage 92.44% 92.41% -0.03%
==========================================
Files 134 130 -4
Lines 7939 7474 -465
Branches 2650 2650
==========================================
- Hits 7339 6907 -432
+ Misses 600 567 -33 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
🔥 Firebase Hosting previews
Channels expire 7 days after their last deploy. |
…ched Review of #495 (F1): the gate read a request the suite REFUSED as one that never arrived. The suite's dispatcher logs "Incoming HTTP request to <path>" before the module runs (TestDispatcher). The "Authorization endpoint" block is only written once the module moves to RUNNING. That throws "Illegal test state change" for a FINISHED/INTERRUPTED instance (AbstractTestModule), and OIDCCClientTestDiscoveryIssuerMismatch throws before the block. So a refused request left no block, and the gate would have rerun past a real verdict. * The gate now requires the instance to still be WAITING, read from GET api/info/{id} before anything is discarded. That is the stall's signature. An instance with a verdict is never rerun, whatever its log says. The log is only read for a WAITING instance. * An "Incoming HTTP request to …/authorize" line counts as arrived, as well as the Authorization endpoint block. * The RERUN note names the discarded instance's status and result. * Unit tests cover the refused-request log shape, FINISHED/FAILED, FINISHED/PASSED and INTERRUPTED/FAILED first attempts, an unknown status, query strings, and other endpoints. Mutation-checked: dropping the WAITING condition or the arrival-line match each fails a test. * The doc comments (api.dart, tests.yaml) now state the real conditions, and say the gate applies on every platform (A1). Also from the review: * F2: the job comment no longer says the boot overlaps Set up Environment. * A2: CONFORMANCE_FORCE_STALL_MODULE (workflow_dispatch input ios_force_stall_module, empty by default) makes the first attempt of one module never open its browser, once per process, to prove the rerun path end to end in CI. * A3: ios_ci_metrics.py closes a login on the RERUN line too, so the discarded attempt's stall still counts. * A4: the collect step tolerates a missing pids file. A dry run under bash -eo pipefail with xcrun stubbed and no pids file exits 0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016g5xVp7yzwhi7f9tLwcNBj
No conflicts. upload-coverage's needs: list (main's split unit_* jobs plus android/ios/web/macos/linux/windows) is unchanged. Its "*-integration-coverage" download pattern still picks up the iOS legs' ios-<shard>-integration-coverage artifacts and not the new ios-<shard>-diagnostics ones. actionlint is clean. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016g5xVp7yzwhi7f9tLwcNBj
On GitHub's arm64 macOS runner (3 vCPU, 7 GB) the simulator swaps constantly, and the system browser sometimes takes 30–150 s to start or never starts. The conformance suite then never receives the authorization request and the module stays WAITING. Screen recording during a stall shows a blank login sheet and no dialog. Lighter host settings, sharding and the Intel runner (2–3× slower, hit the 60-min timeout) were measured across runs and don't remove it.
[ios-metrics]lines. Heavy capture stays as an opt-in dispatch input.Verified on run 37470405361 (both iOS shards green); run 37466674160 without the rerun failed on exactly this stall.
Refs #467
🤖 Generated with Claude Code
https://claude.ai/code/session_016g5xVp7yzwhi7f9tLwcNBj
Summary by CodeRabbit