Repository navigation
fix: alert on stalled rollout-checker CronJobs - #7255
Conversation
Add DocsRolloutCheckerStalled and DocsPrRolloutCheckerStalled alerts to cluster-objects/prometheusrule.yaml, backed by kube-state-metrics' kube_cronjob_status_last_successful_time. Neither nextra-rollout-checker nor nextra-pr-rollout-checker (cluster-objects/job.yaml, pr-job.yaml) had any failure alerting: both are concurrencyPolicy: Forbid with failedJobsHistoryLimit: 1 and no retry beyond the next scheduled tick, so a stuck/failing checker silently stops deploying new images with no signal. Adds runbooks/rollout-checker-failure.md and a runbooks/README.md index entry, and extends src/__tests__/cluster-objects-metrics-consistency.test.ts to recognize kube-state-metrics' metric/labels as a legitimate external source alongside the existing Prometheus built-ins. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: operations <operations@hive.kubestellar.io>
|
[APPROVALNOTIFIER] This PR is NOT APPROVED This pull-request has been approved by: The full list of commands accepted by this bot can be found here. DetailsNeeds approval from an approver in each of these files:Approvers can indicate their approval by writing |
✅ Deploy Preview for kubestellar-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
|
Hi @hivecommons-hive[bot]. Thanks for your PR. I'm waiting for a kubestellar member to verify that this patch is reasonable to test. If it is, they should reply with Once the patch is verified, the new status will be reflected by the I understand the commands that are listed here. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
|
Thank you for your contribution! Your PR has been merged. Check out what's new:
Stay connected: Slack #kubestellar-dev | Multi-Cluster Survey |
Finding
cluster-objects/prometheusrule.yamlalerts on the docs site's own/api/metricsoutput (error rate, latency, scrape-target-down), but theautomated deploy pipeline itself had no alerting:
nextra-rollout-checker(
cluster-objects/job.yaml, every 2m) andnextra-pr-rollout-checker(
cluster-objects/pr-job.yaml, every 1m) poll OCIR and roll out newimages with
set -e,concurrencyPolicy: Forbid,backoffLimit: 0, andonly
failedJobsHistoryLimit: 1— a stuck/failing checker (expiredoci-config-secret, an RBAC regression onnextra-rollout-sa, an OCI APIoutage) silently stops deploying new images with no signal.
Closes #7254.
Change
DocsRolloutCheckerStalled/DocsPrRolloutCheckerStalledalertsto
cluster-objects/prometheusrule.yaml, driven by kube-state-metrics'kube_cronjob_status_last_successful_time(a near-universal companionto a Prometheus Operator install — no new exporter is added, consistent
with every other resource in
cluster-objects/).runbooks/rollout-checker-failure.md(diagnose viakubectl get cronjob/logs, checkoci-config-secretandnextra-rollout-saRBAC)and a
runbooks/README.mdindex entry.src/__tests__/cluster-objects-metrics-consistency.test.tstorecognize kube-state-metrics' metric/labels as a legitimate external
source, alongside the existing Prometheus built-in (
up).No existing alert, SLO, or probe is weakened — this only adds two new
alerts and a runbook.
npx vitest run src/__tests__/cluster-objects-metrics-consistency.test.tspasses (4/4).— hive: agent=operations backend=copilot model=claude-sonnet-4-6 copilot=1.0.88