Parent epic: #45
Related proposal: #84
Release allocation
IncidentRelay 2.4 — lifecycle resilience
- Event Orchestration
set_grouping.reopen_window_seconds; missing/zero = disabled;
- new child Alert on reopen; resolved child Alerts remain terminal;
- one eligible resolved AlertGroup reused inside the window; new group outside it;
- route/team/service/grouping boundaries and resolve/reopen concurrency;
- linked Incident reopen;
- recovery notification hysteresis and cancellation on quick reopen;
- Explain/timeline/audit/UI coverage.
IncidentRelay 2.7 — stale-signal follow-up
- optional stale-signal policy, disabled by default;
- async scheduler/worker evaluation; inactivity alone is not recovery.
IncidentRelay 2.9 — analytics
This workstream is no longer a 2.3 release blocker.
Architectural rules
closed is not added to the technical AlertGroup state machine.
closed remains a terminal operational state of first-class Incident.
- Resolved child Alerts are never mutated back to
firing.
- Reopening an AlertGroup creates a new child Alert occurrence.
- Event Orchestration controls reopen eligibility; the default is disabled.
set_grouping.window_seconds and reopen_window_seconds are separate
concepts and must not share semantics.
- Absence of webhook traffic is not proof of recovery.
2.4 scope
- Extend Event Orchestration validation, execution, simulator and Explain
output for reopen_window_seconds.
- Carry the effective reopen value through orchestration runtime into alert
lifecycle processing.
- Find an eligible recently resolved AlertGroup using the effective grouping
identity and security/routing boundaries.
- Reopen the group atomically and create a new child Alert occurrence.
- Preserve old resolved child Alert rows and timestamps.
- Reset/restart only lifecycle state that the existing firing/reopen rules
require.
- Reconcile linked first-class Incident lifecycle without overwriting
AlertGroup technical state.
- Add API/OpenAPI, audit, timeline, localization and documentation coverage.
- Clarify/fix the existing
set_grouping.window_seconds contract separately:
it must not be repurposed as the resolved-group reopen window.
Acceptance criteria
Out of scope for the 2.4 subset
- automatic Incident closure;
- global always-on reopen behavior;
- treating
set_grouping.window_seconds as the reopen window;
- stale-firing auto-resolution;
- flapping analytics dashboard.
Dependencies
Parent epic: #45
Related proposal: #84
Release allocation
IncidentRelay 2.4 — lifecycle resilience
set_grouping.reopen_window_seconds; missing/zero = disabled;IncidentRelay 2.7 — stale-signal follow-up
IncidentRelay 2.9 — analytics
This workstream is no longer a 2.3 release blocker.
Architectural rules
closedis not added to the technical AlertGroup state machine.closedremains a terminal operational state of first-classIncident.firing.set_grouping.window_secondsandreopen_window_secondsare separateconcepts and must not share semantics.
2.4 scope
output for
reopen_window_seconds.lifecycle processing.
identity and security/routing boundaries.
require.
AlertGroup technical state.
set_grouping.window_secondscontract separately:it must not be repurposed as the resolved-group reopen window.
Acceptance criteria
reopen_window_seconds, current behavior isunchanged and a firing signal after resolution creates a new AlertGroup.
reopen_window_seconds = 0explicitly disables resolved-group reuse.eligible resolved AlertGroup.
child Alert back to firing.
reopened groups or duplicate active occurrences.
state.
closedstate is introduced.configuration.
Out of scope for the 2.4 subset
set_grouping.window_secondsas the reopen window;Dependencies