Parent epic: #45
Release allocation
Target release: IncidentRelay 2.9 — Analytics, Postmortems and final cleanup
Analytics follows the stable 2.3-2.8 domain/lifecycle work. Reports distinguish AlertGroup technical metrics from Incident operational metrics, including #84/#85 reopen/flapping/stale-signal metrics.
Architectural direction
This issue follows the approved Incident Management v2 separation:
Alert -> AlertGroup -> Incident
-
AlertGroup owns technical deduplication, source status, routing, notifications and escalation.
-
Incident owns operational lifecycle, priority, commander, responders, stakeholders, affected services, root cause, resolution and ITSM references.
-
Current technical behavior moves from /api/incidents to /api/alert-groups.
-
/api/incidents is replaced directly and exposes only first-class Incidents.
-
No /api/managed-incidents, compatibility alias or intermediate API is introduced.
-
Manual AlertGroup creation explicitly creates an AlertGroup and initial child Alert.
-
Manual Incident creation creates only a first-class Incident.
Architecture document: docs/architecture/incident-management-v2.md
Goal
Provide separate, correctly defined analytics for technical AlertGroups and operational Incidents.
Scope
-
AlertGroup volume, firing duration, acknowledgement time, notification performance and escalation performance.
-
Classification distribution, false-positive rate, duplicate rate and known-issue rate.
-
Incident volume by owning team, affected service and P1-P5 priority.
-
Time to declaration, investigation, identification, monitoring, resolution and closure.
-
Incident reopen, merge, split, linked-signal and affected-service counts.
-
Root-cause and resolution reporting with privacy-aware access controls.
-
External-ticket coverage, synchronization failures and retry outcomes.
-
CSV/API export with documented UTC and lifecycle definitions.
Acceptance criteria
Flapping/reopen analytics from #84
Lifecycle implementation is tracked in #85.
Add clearly defined AlertGroup metrics for:
-
reopen count;
-
new child occurrences created by resolved-group reuse;
-
flapping occurrence count;
-
time from resolution to reopen;
-
percentage of resolved groups reopened within the configured window;
-
recovery notifications suppressed by hysteresis, when added;
-
stale-signal transitions, when added.
These are AlertGroup/Alert occurrence metrics and must not be counted as
Incident close/reopen metrics unless a first-class Incident transition also
occurred.
Boundary with IncidentRelay 2.10 handling SLA
Tracked follow-up: #96.
2.9 analytics should continue to expose exact existing lifecycle timestamps and
avoid overloading generic MTTR terminology. Alert handling SLA/deadline metrics
are introduced only when the 2.10 handling model exists, so 2.9 does not need
placeholder handling fields or inferred deadlines.
Dependencies
Parent epic: #45
Release allocation
Target release: IncidentRelay 2.9 — Analytics, Postmortems and final cleanup
Analytics follows the stable 2.3-2.8 domain/lifecycle work. Reports distinguish AlertGroup technical metrics from Incident operational metrics, including #84/#85 reopen/flapping/stale-signal metrics.
Architectural direction
This issue follows the approved Incident Management v2 separation:
AlertGroupowns technical deduplication, source status, routing, notifications and escalation.Incidentowns operational lifecycle, priority, commander, responders, stakeholders, affected services, root cause, resolution and ITSM references.Current technical behavior moves from
/api/incidentsto/api/alert-groups./api/incidentsis replaced directly and exposes only first-class Incidents.No
/api/managed-incidents, compatibility alias or intermediate API is introduced.Manual AlertGroup creation explicitly creates an AlertGroup and initial child Alert.
Manual Incident creation creates only a first-class Incident.
Architecture document:
docs/architecture/incident-management-v2.mdGoal
Provide separate, correctly defined analytics for technical AlertGroups and operational Incidents.
Scope
AlertGroup volume, firing duration, acknowledgement time, notification performance and escalation performance.
Classification distribution, false-positive rate, duplicate rate and known-issue rate.
Incident volume by owning team, affected service and P1-P5 priority.
Time to declaration, investigation, identification, monitoring, resolution and closure.
Incident reopen, merge, split, linked-signal and affected-service counts.
Root-cause and resolution reporting with privacy-aware access controls.
External-ticket coverage, synchronization failures and retry outcomes.
CSV/API export with documented UTC and lifecycle definitions.
Acceptance criteria
Incident counts are never calculated by counting all AlertGroups.
Every metric states whether it uses AlertGroup or Incident timestamps.
Reports respect backend team/group RBAC.
Resolved AlertGroups and closed Incidents are not treated as equivalent events.
Exports match UI values and use stable documented field names.
Large date ranges use bounded, indexed and tested queries.
Migration-created Incidents are identified consistently in historical reports.
Flapping/reopen analytics from #84
Lifecycle implementation is tracked in #85.
Add clearly defined AlertGroup metrics for:
reopen count;
new child occurrences created by resolved-group reuse;
flapping occurrence count;
time from resolution to reopen;
percentage of resolved groups reopened within the configured window;
recovery notifications suppressed by hysteresis, when added;
stale-signal transitions, when added.
These are AlertGroup/Alert occurrence metrics and must not be counted as
Incident close/reopen metrics unless a first-class Incident transition also
occurred.
Boundary with IncidentRelay 2.10 handling SLA
Tracked follow-up: #96.
2.9 analytics should continue to expose exact existing lifecycle timestamps and
avoid overloading generic MTTR terminology. Alert handling SLA/deadline metrics
are introduced only when the 2.10 handling model exists, so 2.9 does not need
placeholder handling fields or inferred deadlines.
Dependencies
Incident Management v2: AlertGroup classification and review policy #46
Incident Management v2: Incident domain model and legacy data migration #47
Incident Management v2: Incident links, merge, split and duplicate handling #50
Incident Management v2: External references and generic ITSM framework #52