From 91851d822486b8800ba351c2fb31bce82b2ced5a Mon Sep 17 00:00:00 2001 From: Scot Wells Date: Fri, 25 Sep 2026 18:15:02 -0500 Subject: [PATCH] fix: Scope processor health alerts to the processor The processor's NATS-connection and active-policy gauges are registered in the shared metrics registry of the single activity binary, so every subcommand serves them. The edge event exporter never sets them and reports both at zero, so once the edge exporter is scraped it raises a critical ActivityProcessorNATSDisconnected and an ActivityProcessorNoPolicies warning against a healthy system. Key changes: - Select job="activity-processor" in both rules, matching the processor dashboards and the exporter rules that already select their own job - Leave the rate-based processor rules alone: an idle zero counter cannot satisfy them Changing an expression resets the rule's alert state, so an alert that is firing when this deploys waits out its "for" again. --- config/components/observability/alerts/activity-alerts.yaml | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/config/components/observability/alerts/activity-alerts.yaml b/config/components/observability/alerts/activity-alerts.yaml index 7dc6dc00..f4b0014d 100644 --- a/config/components/observability/alerts/activity-alerts.yaml +++ b/config/components/observability/alerts/activity-alerts.yaml @@ -216,7 +216,7 @@ spec: # NATS Connection Health - alert: ActivityProcessorNATSDisconnected - expr: activity_processor_nats_connection_status == 0 + expr: activity_processor_nats_connection_status{job="activity-processor"} == 0 for: 2m labels: severity: critical @@ -263,7 +263,7 @@ spec: # Policy Cache Health - alert: ActivityProcessorNoPolicies - expr: activity_processor_active_policies == 0 + expr: activity_processor_active_policies{job="activity-processor"} == 0 for: 15m labels: severity: warning