The single-file stdlib server became the limit: no way to track what had been done about an alert, no accounts, and a UI that had to be hand-rolled in template strings. This restructures it into something deployable. Backend (FastAPI) - app/ holds config, database, auth, delivery and the routers; triagelib keeps the triage engine unchanged, so the validated screening and runbook logic is untouched. - Cases persist per alert fingerprint with a status workflow (investigating, customer contacted, escalated to Infra, waiting, remediated, resolved, won't fix, false positive), an assignee, notes and an append-only history. An alert that stops and re-fires lands back on the same case and counts as a reopen. - Suppression rules move from a JSON file into the database. Auth - Signed session cookies over PBKDF2 local accounts, plus an OIDC flow ready for Authentik: users are created on first login and admin follows a group claim. Local login can be switched off entirely once SSO is live. Zendesk and Jira - Delivery is now implemented, behind three gates: the integration must be configured, its feature flag on, and CX_FEATURE_SEND_ENABLED on. A demo instance leaves the last off and cannot mail anyone. Both search before creating, so re-diagnosing an alert updates one ticket rather than opening several, and a rolling daily cap stops a loop mailing everybody. Deployment - Multi-stage Dockerfile builds the bundle and serves it from the API origin. - docker-compose for local and single-host use; Gitea Actions runs the tests, builds the image and renders deploy/k8s with envsubst. Two fixes found while testing: assigning a case returned a null assignee, and add_event could leave an already-loaded history collection stale. Known gap: the engine reaches OpenStack via `docker exec <region>-osc`, which does not work in a pod without the CX-Tools containers alongside it. docs/DEPLOYMENT.md sets out the three ways to close that. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
3.6 KiB
Deployment
Local
cp .env.example .env # set CX_SECRET_KEY and CX_BOOTSTRAP_ADMIN_PASSWORD
docker compose up --build
http://localhost:8080, sign in with the bootstrap admin.
Without Docker:
cd backend && pip install -r requirements.txt
uvicorn app.main:app --reload --port 8080 # terminal 1
cd frontend && npm install && npm run dev # terminal 2 -> :5173, proxies /api
Kubernetes via Gitea Actions
.gitea/workflows/ci.yaml runs tests → builds the image → renders
deploy/k8s/*.yaml with envsubst → applies them.
Repository variables (not secret)
| Variable | Example | Meaning |
|---|---|---|
CX_DOMAIN |
cx-triage.ngbackend.cloud |
Ingress host; CX_BASE_URL derives from it |
K8S_NAMESPACE |
cx-triage |
Target namespace |
REGISTRY |
git.ngbackend.cloud |
Image registry |
IMAGE_NAME |
parham.monfared/cx-ui |
Image repository |
CX_PROMETHEUS_BASE |
http://10.11.254.250:9090 |
Alert source |
CX_OIDC_ENABLED |
true |
Authentik on |
CX_OIDC_ISSUER |
https://sso…/application/o/cx-triage/ |
Discovery base |
CX_OIDC_ADMIN_GROUP |
cx-triage-admins |
Group granting admin |
CX_FEATURE_SEND_ENABLED |
false |
Master send switch |
CX_FEATURE_ZENDESK / CX_FEATURE_JIRA |
false |
Per-integration flags |
CX_FEATURE_LINKAGE_SCAN |
true |
Expensive scan on/off |
CX_SEND_DAILY_CAP |
25 |
Rolling 24h send limit |
CX_JIRA_PROJECT / CX_JIRA_BASE |
INFRA |
Jira target |
INGRESS_CLASS / TLS_ISSUER |
nginx / letsencrypt-prod |
Ingress wiring |
REPLICAS |
1 |
See the caveat below |
Repository secrets
KUBECONFIG (base64), REGISTRY_USERNAME, REGISTRY_TOKEN, CX_SECRET_KEY,
CX_DATABASE_URL, CX_OIDC_CLIENT_ID, CX_OIDC_CLIENT_SECRET,
CX_ZENDESK_SUBDOMAIN, CX_ZENDESK_EMAIL, CX_ZENDESK_TOKEN,
CX_JIRA_EMAIL, CX_JIRA_TOKEN, CX_BOOTSTRAP_ADMIN_PASSWORD.
Authentik
- Applications → Providers → Create → OAuth2/OpenID Provider
- Client type: Confidential
- Redirect URI:
https://<CX_DOMAIN>/api/auth/oidc/callback - Scopes:
openid,email,profile
- Create the Application and bind the provider.
- Copy the client ID/secret into the Gitea secrets above.
- Create a group
cx-triage-admins; its members get admin rights. - Set
CX_OIDC_ISSUERto the provider's OpenID configuration base URL — the app appends/.well-known/openid-configuration.
Users are created on first login. CX_AUTH_LOCAL_ENABLED=false in the cluster
config turns off password login entirely once SSO works.
Two things to plan for
Reaching OpenStack. The diagnosis engine shells out to
docker exec <region>-osc openstack …, and the internal Prometheus is only
reachable from inside those containers. That works on a laptop with CX-Tools
running; it does not work in a pod by default. Options, cheapest first:
- Run the app on a host that already has the CX-Tools containers, mounting the
Docker socket (what
docker-compose.ymldoes). - Run the
*-osccontainers as sidecars in the pod. - Replace
cxbridge.os_jsonwith direct authenticated OpenStack API calls and give the pod a network route. Cleanest, most work.
Until one of those is in place, a cluster deployment can read Prometheus (if routable) but per-alert diagnosis will fail. That is a real gap, not an oversight.
REPLICAS should stay at 1 for now. The Prometheus caches, the background
triage jobs and the linkage scan are per-process, so a second replica would
duplicate the work and serve inconsistent job IDs. Moving jobs into the database
or a queue is what unlocks scaling out.