Alerting
Metric alerts are evaluated from Prometheus/Mimir data in Grafana. Log, trap, event, and CDR/CMR alerts are evaluated as scheduled searches in Splunk. Selected notifications are sent to EMS/KeepHQ and the configured response destinations.
Architecture
graph LR
A[Mimir Metrics] --> B[Grafana Alert Rules]
C[Splunk Events] --> D[Splunk Alerts]
B --> E[Contact Points]
D --> E
E --> F[EMS / KeepHQ]
F --> G[Webex]
F --> H[PagerDuty for Selected Critical Alerts]
F --> I[Other Configured Destinations]
Alert evaluation and notification delivery are separate stages. During troubleshooting, confirm the source rule fired, a notification policy/contact point selected it, EMS accepted it, and the final destination delivered it.
Components
| Component | Purpose |
|---|---|
| Grafana Alerts | Metric-based service, resource, and scrape-health alerting |
| Splunk Alerts | Log, trap, event, and CDR/CMR pattern alerting |
| EMS / KeepHQ | Notification intake, normalization, routing, and delivery |
Not every device category currently has complete dashboard, service-alert, and scrape-down coverage. The individual device pages identify the gaps recorded in the audit repository.