Skip to content

Incident Management v2: Incident and alert-quality analytics #54

Description

@Aidaho12

Parent epic: #45

Release allocation

Target release: IncidentRelay 2.9 — Analytics, Postmortems and final cleanup

Analytics follows the stable 2.3-2.8 domain/lifecycle work. Reports distinguish AlertGroup technical metrics from Incident operational metrics, including #84/#85 reopen/flapping/stale-signal metrics.

Architectural direction

This issue follows the approved Incident Management v2 separation:


Alert -> AlertGroup -> Incident

  • AlertGroup owns technical deduplication, source status, routing, notifications and escalation.

  • Incident owns operational lifecycle, priority, commander, responders, stakeholders, affected services, root cause, resolution and ITSM references.

  • Current technical behavior moves from /api/incidents to /api/alert-groups.

  • /api/incidents is replaced directly and exposes only first-class Incidents.

  • No /api/managed-incidents, compatibility alias or intermediate API is introduced.

  • Manual AlertGroup creation explicitly creates an AlertGroup and initial child Alert.

  • Manual Incident creation creates only a first-class Incident.

Architecture document: docs/architecture/incident-management-v2.md

Goal

Provide separate, correctly defined analytics for technical AlertGroups and operational Incidents.

Scope

  • AlertGroup volume, firing duration, acknowledgement time, notification performance and escalation performance.

  • Classification distribution, false-positive rate, duplicate rate and known-issue rate.

  • Incident volume by owning team, affected service and P1-P5 priority.

  • Time to declaration, investigation, identification, monitoring, resolution and closure.

  • Incident reopen, merge, split, linked-signal and affected-service counts.

  • Root-cause and resolution reporting with privacy-aware access controls.

  • External-ticket coverage, synchronization failures and retry outcomes.

  • CSV/API export with documented UTC and lifecycle definitions.

Acceptance criteria

  • Incident counts are never calculated by counting all AlertGroups.

  • Every metric states whether it uses AlertGroup or Incident timestamps.

  • Reports respect backend team/group RBAC.

  • Resolved AlertGroups and closed Incidents are not treated as equivalent events.

  • Exports match UI values and use stable documented field names.

  • Large date ranges use bounded, indexed and tested queries.

  • Migration-created Incidents are identified consistently in historical reports.

Flapping/reopen analytics from #84

Lifecycle implementation is tracked in #85.

Add clearly defined AlertGroup metrics for:

  • reopen count;

  • new child occurrences created by resolved-group reuse;

  • flapping occurrence count;

  • time from resolution to reopen;

  • percentage of resolved groups reopened within the configured window;

  • recovery notifications suppressed by hysteresis, when added;

  • stale-signal transitions, when added.

These are AlertGroup/Alert occurrence metrics and must not be counted as

Incident close/reopen metrics unless a first-class Incident transition also

occurred.

Boundary with IncidentRelay 2.10 handling SLA

Tracked follow-up: #96.

2.9 analytics should continue to expose exact existing lifecycle timestamps and
avoid overloading generic MTTR terminology. Alert handling SLA/deadline metrics
are introduced only when the 2.10 handling model exists, so 2.9 does not need
placeholder handling fields or inferred deadlines.

Dependencies

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    incident-management-v2Incident Management v2 architecture and delivery

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions