Summary
The automation docs/INCIDENT-RESPONSE.md's "Main Branch Build Recovery SLA"
depends on still does not exist, re-confirmed this session:
```bash
$ test -f .github/on-call-schedule.yml && echo exists || echo MISSING
MISSING
$ gh label list --repo kubestellar/console --search main-broken
main-legacy Main legacy related #82FCBE
$ grep -rln "main-broken|kubestellar-dev|kubestellar-maintainers" .github/workflows/*.yml
(no output)
```
This is the same gap originally tracked by #23265 and re-confirmed by #23367 — both
are now closed (the latter only fixed dead-tracker references in
`docs/runbooks/incident-response-automation-missing.md` and
`docs/runbooks/llmd-guide-e2e-silent-failure.md` via PR #23368, not the automation
itself). Filing a fresh tracker so the runbook has a live issue to cite, per the
dead-tracker pattern already documented in #23367.
Gap detail (unchanged from #23265)
.github/on-call-schedule.yml — referenced in docs/INCIDENT-RESPONSE.md as the
Build Sheriff rotation source of truth, still marked "(to be created)", does not
exist.
- The
main-broken label — docs/INCIDENT-RESPONSE.md says this is auto-applied
to the last merged PR when main CI fails. It does not exist in the repo's label
list.
- No workflow references
main-broken, kubestellar-dev, or
kubestellar-maintainers — the doc's "Detection (Automated)" section describes
a Slack-post-and-label step that isn't wired to anything.
Why this matters
Without these, a broken main branch produces only a red run in the Actions tab —
no automated notification, no label, and no named owner for the 4-hour recovery
clock the rest of the incident-response playbook (escalation matrix, circuit
breaker, postmortem step) is built on.
Why operations can't fix this directly
- The workflow-step portion of the fix would touch
.github/workflows/**, which
this agent's App token cannot push (no workflows permission) — needs a human
or an ISSUES_PRS_MERGE-tier agent.
- Naming an actual on-call rotation for
.github/on-call-schedule.yml is a
personnel/process decision this agent cannot make.
- Creating the
main-broken label itself does not require a workflow change and
could be done by any maintainer/agent with label-create access.
Proposed fix
- A maintainer creates
.github/on-call-schedule.yml naming the actual current
Build Sheriff rotation.
- Create the
main-broken label.
- Add a workflow step, gated on the main-branch build/test jobs with
if: failure(), that applies the label and opens/updates an incident-tracking
issue — reusing the same create-or-update-issue pattern already implemented in
.github/workflows/workflow-failure-issue.yml.
- Alternatively, edit
docs/INCIDENT-RESPONSE.md §"Detection (Automated)" to stop
describing automation that doesn't exist, so the doc doesn't overstate current
coverage until the wiring lands.
docs/runbooks/incident-response-automation-missing.md is being updated in the
companion PR to cite this issue instead of the now-closed #23367.
🐝 Hive Agent: operations | Instance: hosted-kubestellar-console-4vkt | SHA: 30fce13
— hive: agent=operations backend=copilot model=claude-sonnet-4-6 copilot=1.0.78
Summary
The automation
docs/INCIDENT-RESPONSE.md's "Main Branch Build Recovery SLA"depends on still does not exist, re-confirmed this session:
```bash
$ test -f .github/on-call-schedule.yml && echo exists || echo MISSING
MISSING
$ gh label list --repo kubestellar/console --search main-broken
main-legacy Main legacy related #82FCBE
$ grep -rln "main-broken|kubestellar-dev|kubestellar-maintainers" .github/workflows/*.yml
(no output)
```
This is the same gap originally tracked by #23265 and re-confirmed by #23367 — both
are now closed (the latter only fixed dead-tracker references in
`docs/runbooks/incident-response-automation-missing.md` and
`docs/runbooks/llmd-guide-e2e-silent-failure.md` via PR #23368, not the automation
itself). Filing a fresh tracker so the runbook has a live issue to cite, per the
dead-tracker pattern already documented in #23367.
Gap detail (unchanged from #23265)
.github/on-call-schedule.yml— referenced indocs/INCIDENT-RESPONSE.mdas theBuild Sheriff rotation source of truth, still marked "(to be created)", does not
exist.
main-brokenlabel —docs/INCIDENT-RESPONSE.mdsays this is auto-appliedto the last merged PR when main CI fails. It does not exist in the repo's label
list.
main-broken,kubestellar-dev, orkubestellar-maintainers— the doc's "Detection (Automated)" section describesa Slack-post-and-label step that isn't wired to anything.
Why this matters
Without these, a broken main branch produces only a red run in the Actions tab —
no automated notification, no label, and no named owner for the 4-hour recovery
clock the rest of the incident-response playbook (escalation matrix, circuit
breaker, postmortem step) is built on.
Why operations can't fix this directly
.github/workflows/**, whichthis agent's App token cannot push (no
workflowspermission) — needs a humanor an ISSUES_PRS_MERGE-tier agent.
.github/on-call-schedule.ymlis apersonnel/process decision this agent cannot make.
main-brokenlabel itself does not require a workflow change andcould be done by any maintainer/agent with label-create access.
Proposed fix
.github/on-call-schedule.ymlnaming the actual currentBuild Sheriff rotation.
main-brokenlabel.if: failure(), that applies the label and opens/updates an incident-trackingissue — reusing the same create-or-update-issue pattern already implemented in
.github/workflows/workflow-failure-issue.yml.docs/INCIDENT-RESPONSE.md§"Detection (Automated)" to stopdescribing automation that doesn't exist, so the doc doesn't overstate current
coverage until the wiring lands.
docs/runbooks/incident-response-automation-missing.mdis being updated in thecompanion PR to cite this issue instead of the now-closed #23367.
🐝 Hive Agent:
operations| Instance:hosted-kubestellar-console-4vkt| SHA:30fce13— hive: agent=operations backend=copilot model=claude-sonnet-4-6 copilot=1.0.78