# From "here's what to do" to "do it" ## 1. Validation against live data Every recommendation was checked against OpenStack directly, not against the app's own view. Results: | Case | App said | Ground truth | Verdict | |---|---|---|---| | **Rogue VM `CA1-ESC8-057`** | "No mismatch — go check InfraInsight for HIBERNATED VMs carrying a host" | OpenStack 7 GPUs across 5 VMs; Infrahub 7 GPUs across the same 5 VMs. Perfect agreement. No hibernated VMs on the host at all. | **Wrong** | | **Rogue VM, 10 hosts sampled** | 50 alerts "needs action" | OpenStack and Infrahub agreed **exactly** on all 10 | **Wrong** | | `SHUTOFF` sn56-week6-transfer-e1 | Billing notice to `shettyatulya@gmail.com` | OpenStack `SHUTOFF`, no fault | Correct | | `DELETING` hs-gaussian-splatting | Intent verified from events; delete in OpenStack | OpenStack `ACTIVE`, delete requested 17:00 | Correct | | `DELETING` c3-pool-1785947188798 | Intent verified; delete in OpenStack | OpenStack `ERROR`, host `None`, "No valid host was found" | Correct action, **incomplete** — it never built, so the customer is owed the creation-failure notice too | | `ERROR` hyperstack-minion-lydhjev | `no_valid_host` fault → never ACTIVE → insufficient-stock template → Lars Vagnes | Fault is exactly "No valid host was found. There are not enough hosts available."; host `None` | Correct — but was **hidden as "chronic"** | ### The Rogue VM rule is broken at source ```promql sum by (instance) (In_Use_Gpus) - sum by (instance) (Resources{status=~"ACTIVE|SHUTOFF|PRE_ACTIVE"}) >= 1 ``` `In_Use_Gpus` equals `Total_Gpus` — the **physical** GPU count — on **71 of the 75** firing hosts. So the expression reduces to *physical GPUs minus allocated GPUs*, i.e. **"this host has at least one free GPU."** That is spare capacity, not a rogue VM. `CA1-ESC8-057` is the clean example: 8 physical, 7 allocated across 5 VMs that both systems agree on, so it fires with a gap of 1. Two consequences: - 65 of the 72 rogue-VM alerts are now classified **invalid — alert rule defect** and screened out, with the reason stated. - The 2 that survive are genuine: `CA1-ESC812-252` (In_Use 2, one 1-GPU VM) and `CA1-ESC812-289` (In_Use 7, 5 GPUs allocated). Both now read *"N GPU(s) in use belong to no instance on either side"* and route to Infrastructure, because no CX action can fix a leaked hypervisor allocation. **This needs fixing in the Prometheus rule, not just filtered here.** Whoever owns `infrahub-rules.yml` should either repair `In_Use_Gpus` or rewrite the expression. ### "Chronic" was hiding overdue customer contact 13 ERROR alerts have held for 6–7 days. The ERROR runbook says to contact the customer if a stock-failure instance is not deleted **within a day**. Filing those as "chronic, probably already ticketed" was wrong — they are overdue. New `overdue` verdict: for kinds with a runbook time commitment, age makes a case *more* urgent and it can never be demoted to chronic. ### Net effect | | before validation | after | |---|---|---| | To action | 54 | **19** | | of which overdue customer contact | 0 (hidden) | **13** | | Rogue VM noise | 50 "needs action" | 65 flagged as a rule defect | --- ## 2. Integration plan The principle: the app already knows *who* to contact and *what* to say. Wiring delivery turns a 5-minute copy-paste into one reviewed click — without ever sending on its own. ### Phase 1 — Zendesk (the main win) `integrations.build_zendesk()` already produces a complete `POST /api/v2/tickets` body: requester resolved from Infrahub owners, subject, the approved runbook wording with placeholders filled, priority derived from the verdict, tags, and `external_id = cx-triage-` for idempotency. To finish it: 1. **Credentials** — `CX_ZENDESK_SUBDOMAIN`, `CX_ZENDESK_EMAIL`, `CX_ZENDESK_TOKEN` (API token, Basic auth as `email/token:token`). Read from 1Password via the existing CX-Tools loader rather than env vars, so nothing lands on disk. 2. **Search before create** — `GET /api/v2/search?query=external_id:` so a re-diagnosed alert updates the existing ticket instead of opening a duplicate. 3. **Send** — implement `App.send_zendesk`, which currently refuses by design. 4. **Guard rails** (all already scaffolded in the UI): - two-step confirm naming the recipient - never auto-send; no bulk send in v1 - an `--allow-send` startup flag, so a demo instance physically cannot email - append the ticket URL back onto the case and log it to `vmc-audit.log` 5. **Requester matching** — Infrahub gives owner name + email; Zendesk may already have that user. Search by email, fall back to creating the requester inline. ### Phase 2 — Jira `integrations.build_jira()` produces the `POST /rest/api/3/issue` body with the evidence block already assembled. Needs: project key confirmation (`INFRA`?), issue type, and the same search-before-create against a `cx-triage-` label. The GPU-leak escalations above are the immediate use case. ### Phase 3 — closing the loop - **Slack** — react to the alert in `#infrahub-errors` and thread the findings, which the runbook asks for manually today. - **Case state** — persist handled/snoozed/ticketed per fingerprint (SQLite) so the queue reflects work already done and survives a restart. - **Bulk actions** — the 13 overdue ERROR alerts are one org-grouped mail-merge; worth doing only once single-send is trusted. ### What stays manual, deliberately Deleting, shelving and InfraInsight edits stay copy-a-command. The read-only guarantee is what makes this safe to run against production, and the destructive steps are exactly where a wrong verdict would be expensive. --- ## 3. UI Live at `http://127.0.0.1:8765/` (previous version kept at `/classic`). - **Left rail** — one line per case: subject, org, region, how long the condition has actually held. Grouped and collapsible. Verdict filter chips across the top. - **Case card** — verdict as a sentence, then a *picture* of the problem: - state alerts: `Infrahub says X` ≠ `OpenStack says Y` - GPU cases: a segmented allocation bar (green allocated / red unaccounted) - duplicate IPs: one row per claimant - **"Do this"** — the customer's name and address, then the action buttons. Nothing else competes with them. - **Everything else collapsed** — Caveats, Why, Runbook steps, Raw evidence. - **Zendesk drawer** — recipient, subject, editable body, priority, tags, and a Send button that is disabled and labelled *"Preview only. Nothing will be sent."* until credentials exist and sending is explicitly enabled.