Skip to content

Alerting

Metric alerts are evaluated from Prometheus/Mimir data in Grafana. Log, trap, event, and CDR/CMR alerts are evaluated as scheduled searches in Splunk. Selected notifications are sent to EMS/KeepHQ and the configured response destinations.

Architecture

graph LR
    A[Mimir Metrics] --> B[Grafana Alert Rules]
    C[Splunk Events] --> D[Splunk Alerts]
    B --> E[Contact Points]
    D --> E
    E --> F[EMS / KeepHQ]
    F --> G[Webex]
    F --> H[PagerDuty for Selected Critical Alerts]
    F --> I[Other Configured Destinations]

Alert evaluation and notification delivery are separate stages. During troubleshooting, confirm the source rule fired, a notification policy/contact point selected it, EMS accepted it, and the final destination delivered it.

Components

Component Purpose
Grafana Alerts Metric-based service, resource, and scrape-health alerting
Splunk Alerts Log, trap, event, and CDR/CMR pattern alerting
EMS / KeepHQ Notification intake, normalization, routing, and delivery

Not every device category currently has complete dashboard, service-alert, and scrape-down coverage. The individual device pages identify the gaps recorded in the audit repository.