Operation

Legacy Prometheus is rolled out with a standalone deployment layout. This is simple, however, big cardinality of series in Prometheus Head causes high memory utilization which finally leads out-of-memory issue. All metrics are stored in one individual pod which restricts the extension of capability. The approach to resolve OOM is to seperate metrics of platform from application metrics. By the seperation, the metrics governance will be much easier for different tenants.

Below dashboard tells the overall progress to apply this split prometheus solution on production.

Detailed explanation on split Prometheus could be found HERE