Read-only triage for the Infrahub error alerts. Pulls the Prometheus alert
queue, re-checks each alert's condition against live state to separate real
work from noise, diagnoses it using the CX runbooks, and drafts the customer
comms with contacts resolved from Infrahub.
Findings from validating against production:
- "Suspected Rogue VM" fires on spare GPU capacity, not rogue VMs: In_Use_Gpus
equals the physical count on 71 of 75 firing hosts, so the rule reduces to
"this host has a free GPU". Verified against OpenStack on 10 hosts.
- "Exists in Infrahub but does not exist in OpenStack" matches every VM because
openstack_nova_server_status returns no series; excluded as a rule defect.
- Prometheus activeAt is reset several times a day by dips in the Resources
metric, so alert ages are recovered from ALERTS history instead.
Takes ~2,650 firing alerts down to ~20 that need a decision.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>