# CX Triage A small local webapp that takes the Infrahub error alerts out of Prometheus, diagnoses each one using the **unmodified** CX-Tools (`vmc`) collectors, tells you what the runbook says to do next, and — when the next step is contacting the customer — shows the approved wording alongside the customer's contact details. **It is read-only.** It queries Infrahub, OpenStack, InfraInsight and Prometheus. It never changes platform state, never deletes or shelves anything, and never sends a message. Every action it identifies is presented for a human to perform. ## Separating noise from real work Thousands of alerts fire; only a handful are work. Before anything is shown, each alert's condition is **re-checked against current state**, and the verdict is displayed with its reason: | Verdict | Meaning | In the queue? | |---|---|---| | **needs action** | The condition still holds | yes | | **needs action (unverified)** | Couldn't be re-checked — never hidden on a guess | yes | | **already resolved** | Infrahub has moved on / the IP is no longer duplicated / the GPU gap has closed | hidden | | **not yet firing** | Prometheus still has it pending | hidden | | **chronic** | Still true, but firing over 3 days — already ticketed, not new work | hidden | | **low impact** | Still true, but owned by an internal org or a platform-owned node | hidden | Screening only ever demotes an alert on **positive evidence**; anything it can't settle stays in the queue. Hidden alerts are one checkbox away, and any of them can be force-diagnosed with **Diagnose anyway**. ### Alert ages are recovered, not taken from Prometheus Prometheus' own `activeAt` is unreliable here. The Infrahub `Resources` metric drops most of its series for ~5 minutes several times a day (4 dips in the last 24h observed; one took it from ~4,370 series to 1,359). Every alert alive during a dip resolves and re-fires, so `activeAt` resets on all of them at once — which is why the Prometheus UI shows dozens of unrelated alerts with the *same* age. So the app walks the `ALERTS` series backwards over 7 days instead, bridging gaps under 45 minutes, and reports how long each condition has **actually** held. In practice this is the difference between "40 alerts all 7h old" and "11 that are genuinely new, 33 that have been true for days". Both numbers are shown: the recovered duration, with Prometheus' value in a tooltip when they disagree. The app detects these dips and warns about them, since they also mean any alert with a long `for:` may never reach firing state. The re-check is cheap on purpose: it reads the same Prometheus series the rules are built from — `Resources`, `In_Use_Gpus`, `Total_Gpus` — in one bulk snapshot for the entire queue, rather than an Infrahub and OpenStack call per alert. Only alerts you actually open cost a CX-Tools query. On live data this takes **~2,670 firing alerts down to ~11** that need a decision. Caches are warmed at startup (~20s, mostly the 7-day history read), so page loads are instant afterwards. They refresh on a 30s/60s/5min cadence. **Excluded outright:** `Exists in Infrahub but does not exist in OpenStack`. It is built as `Resources unless on(openstack_id) openstack_nova_server_status`, and that second metric is currently returning **zero series** — so nothing gets excluded by the `unless` and every Infrahub VM alerts. It is a broken exporter, not a queue of work. The app detects this class of failure and shows a banner, because the same gap also means `Suspected Orphan VM` cannot fire at all. ## Two tabs - **CX runbooks** — the alerts below, grouped into collapsible sections in working order (rogue VMs, duplicate IPs, total GPUs, hibernating, creating, shutoff, deleting, error), each showing how long it has been firing, **newest first** so long-running alerts sink to the bottom. - **Infrastructure** — everything else, so it stays out of the triage queue: node-exporter host alerts in their own section, then Ceph, MySQL, Galera, OpenStack services, blackbox. Listed and counted, not diagnosed. Routing is keyed off the **rule file**, not the alert name, because two different rule files both use the group name "Imported Rules". ## What it covers One runbook per alert type, from *Infrahub Errors Remediation*: | Alert | Priority | What the app works out for you | |---|---|---| | Instance in ERROR state | LOW–HIGH | Matches the fault against the runbook fault table; decides whether the VM was ever ACTIVE (which changes both the urgency and the comms template); for the NUMA/PCI fault it sums GPUs on the host to check whether the host is full before you escalate | | Instance in DELETING state | LOW | Confirms the delete request in Infrahub events, and whether the OpenStack server is still there or already gone | | Instance in SHUTOFF state | LOW | Confirms SHUTOFF and drafts the billing-awareness note | | Instance in HIBERNATING state | HIGH | Runs the host signals (Nova state/status, disabled reason, OVS liveness) and escalates when they're bad | | Instance in CREATING state | MEDIUM | Determines whether the VM ever got an OpenStack ID | | Instance in RESTORING state | HIGH | Host signals plus the most recent *failed* OpenStack event to escalate | | Instance in REBOOTING state | HIGH | Confirms `InstanceRebootRequest` and the expected `HARD_REBOOT` state | | Instance in BUILD state | MEDIUM | Distinguishes "large flavor, still transient" from "stuck, escalate" | | Suspected Rogue VM | HIGH | Quantifies the per-host GPU accounting gap the rule actually fires on, then reconciles every instance on the host and maps each mismatch to its row in the Mismatch Remediation table; tempest instances are ignored | | Duplicated IPs | HIGH | Classifies each claimant of the IP as Scenario #1 / #2 / rightful owner, and pulls the cross-environment claimant list from `Resources{floating_ip=...}` | | Problem with Total GPUs | HIGH | Lists which customers are on the affected host | | Suspected Orphan VM | HIGH | Same host reconciliation as Rogue VM (cannot currently fire — see the exporter note above) | | Openstack status=X / Infrahub status!=X | HIGH | Single-VM mismatch, taken through the same remediation table | ### A note on Suspected Rogue VM The rule is not a status comparison — it is `sum by(instance)(In_Use_Gpus) - sum by(instance)(Resources{status=~"ACTIVE|SHUTOFF|PRE_ACTIVE"}) >= 1`, a **per-host GPU accounting gap**. Two different faults produce that gap: 1. instances running on the host that Infrahub has no record of (a true rogue VM), or 2. Infrahub VMs that are ACTIVE but have **no host recorded**, so they are never counted against the host that is actually running them. Prometheus cannot tell these apart, so the app states both and lets the per-instance host reconciliation settle it — a genuine rogue VM shows up as `Infrahub Missing`. It also reports how many unattributed ACTIVE VMs exist platform-wide, because that number alone can be large enough to explain the gaps without any rogue VM existing. ## Requirements Whatever `vmc` already needs, plus nothing: - The CX-Tools checkout (`cxlib/` + `vmc`), unmodified - Python 3 (standard library only — no pip install) - Docker with the `ca1-osc` / `ca2-osc` / `us1-osc` / `no1-osc` containers running - A signed-in 1Password CLI session ## Run ```bash op signin ``` ```bash ./cx-triage ``` It opens . Bound to localhost only. ```bash ./cx-triage --check ``` Runs preflight (CX-Tools located, credentials loaded, containers up, Prometheus reachable) and exits. Useful flags: `--port`, `--prometheus `, `--no-open`. If CX-Tools isn't found automatically, point at it: ```bash CX_TOOLS_PATH=~/scripts/CX-Tools ./cx-triage ``` ## How it reaches things - **CX-Tools** is imported as a library. `cxbridge.py` calls `collect_vm`, `collect_host` and the query helpers — the same code paths as `vmc --json` — and guards every OpenStack subcommand against a read-only allowlist, so a bug here cannot mutate an instance. - **Prometheus** at `10.11.254.250:9090` is on the internal network, which the laptop has no route to (10.11.* leaves via the default gateway). So queries are relayed `docker exec ca1-osc curl ...` — the same trick CX-Tools uses for OpenStack. A direct HTTP transport is tried first, so this still works from a host that does have a route. Override the relay with `CX_PROMETHEUS_RELAY`. ## Working an alert 1. Pick an alert from the queue. Or paste an `ALERTS{...}` line or a Prometheus graph URL into the box. 2. **What the platforms say** — the reconciled Infrahub/OpenStack/InfraInsight facts, with mismatches called out in red. 3. **Next steps** — the remaining runbook steps, each tagged with its owner (CX / Infrastructure team / DevOps). Steps the app has already verified are ticked off, so you can see what's left rather than re-deriving it. 4. **Suggested customer comms** — only when the runbook calls for it. Verbatim approved wording with the instance name (and floating IP) substituted, above the organization and owner contacts CX-Tools resolved. Copy it and send it from HubSpot. 5. **Evidence** — the raw alert labels and the raw CX-Tools output, for pasting into a Slack thread or a Jira ticket. ## Caveats These are real limits, not bugs: - **Production Infrahub only.** CX-Tools queries production. When a Duplicated IPs or Rogue VM alert points at a PreProd/Staging record, the app says so and tells you to check the other environments — it can't query them. The Prometheus `Resources` series does span environments, which is why the Duplicated IPs view uses it. - **No InfraInsight SQL.** The DELETING runbook identifies the requesting *user* via a SQL query. The app confirms the delete request from Infrahub events but cannot name the requester, and says so when it matters. - **Host Health Checks is partial.** The app reports the host signals CX-Tools exposes (Nova state/status, disabled reason, OVS liveness/heartbeat, uptime, aggregates). The rest of that guide — disk, dmesg, GPU checks — is still manual, and the app says which part it did. - **Fault table coverage.** ERROR faults outside the runbook's table produce an explicit "not in the table, escalate to a peer" verdict rather than a guess. - **Chronic/low-impact thresholds are judgement calls**, not runbook rules: 3 days for chronic (`CHRONIC_DAYS` in `screening.py`), and "internal" means an `@nexgencloud.com` owner. Adjust to taste. - **Recovered ages are bounded by a 7-day window** (`TrueAgeIndex.WINDOW_DAYS`). Anything older shows as `7d+`. - **"Chronic" does not mean "ignore".** It means the condition has been true for days, so it is not *new* work. Several ERROR alerts are 6–7 days old; if those have not actually been ticketed, they are a backlog, not noise. ## Layout ```text cx-triage entry point (preflight, then serve) triagelib/ cxbridge.py read-only adapter over cxlib prometheus.py alert source, rule index, bulk state snapshot, relay alerts.py normalization, classification, exclusions, grouping screening.py noise-vs-real verdicts runbooks.py the decision engine comms.py customer comms templates, verbatim from Confluence server.py HTTP API + background triage jobs ui.py the single-page UI tests/test_runbooks.py runbook decisions against fixture payloads tests/test_screening.py screening, exclusion, tab routing, ordering ``` ```bash python3 tests/test_runbooks.py && python3 tests/test_screening.py ``` ## Updating a runbook The decision logic is meant to be edited by whoever owns the runbook: - Fault table → `FAULT_TABLE` in `runbooks.py` - Rogue VM mismatch table → `MISMATCH_TABLE` in `runbooks.py` - Priority / ETTR → `KIND_META` in `alerts.py` - Queue order → `FOCUS_ORDER` in `alerts.py` - Alerts to suppress → `EXCLUDED_ALERTNAMES` in `alerts.py` - Noise rules → `screening.py` (`CHRONIC_DAYS`, `_KIND_SCREENS`) - Customer wording → `_TEMPLATES` in `comms.py` If a wording change lands in Confluence, change it in `comms.py` and nowhere else.