Two real gaps found after testing the live standby platform:
1. modules/kubernetes.py's in-cluster client (load_incluster_config())
can only ever work inside the real k8s pod. On the standby it silently
returned K8S_AVAILABLE=False with zero fallback — Cluster, Containers-
via-sites, and any cluster/app health data was just empty on that page,
not erroring, so it was easy to miss. This holds even if the standby
later runs as a pod on a *local* k3s on that server (per plan for the
monitoring migration) — in-cluster config there would only ever see
that local, unrelated cluster, never the main one.
Fixed by adding an SSH+kubectl remote path to list_pods(), list_
deployments(), get_ingress_info(), list_deployments_for_namespace(),
and get_pod_metrics(): when not RUNNING_ON_MAIN_SERVER, each runs
`kubectl <verb> -o json` on the main server over the tunnel
(main-to-vm-tunnel.service) instead of using the python client, parses
the raw (camelCase) JSON directly rather than trying to force it
through the client-library's snake_case object model, and feeds the
same _pod_health/_deploy_health/_age_str helpers either way. The
in-cluster path for the real main-server pod is untouched.
2. get_system_info() called psutil.* unconditionally with no
RUNNING_ON_MAIN_SERVER awareness at all — on the standby it was
silently reporting the VM's OWN cpu/mem/disk/hostname mislabeled as
"system info" (hostname field literally showed the VM's hostname).
Added _get_main_server_system_info_remote(): SSHes to the main server
and gathers the same stats via vmstat/free/df/procfs (the main
server's bare host has no psutil — that's only in this app's own
container image — so this avoids depending on it being present
remotely).
Verified end-to-end on the live standby: /api/system now reports the
real main server's hostname/cpu/mem/disk/uptime; /api/cluster reports
real counts (19 pods, 16/16 deployments healthy, 3 degraded, plus live
per-pod metrics via the metrics-server raw API); /api/sites shows real
per-app container/role status. All previously silently empty.
modules/sites.py was still a fully Docker-era hardcoded registry:
container names (odoo-clean-odoo-1, frappe-erpnext, nextcloud-app,
mautic-app, n8n-app) that no longer exist post-migration, checked via
`docker inspect` over SSH, plus hardcoded domain/port fields -
nextcloud and mautic even had domain: None, silently falling back to
dead Docker host-port URLs while their real Ingress domains
(next.cloud.nav.ovh, mautics.nav.ovh) sat unused.
Now sources everything live from the cluster:
- Domain + TLS from the actual Ingress object per app (via the new
get_ingress_info() in modules/kubernetes.py), not a static guess.
Odoo/Nextcloud's Ingress lives in `default` (see prior commit for the
RBAC this needed); n8n/mautic/erpnext's lives in their own namespace.
- App/DB/cache/worker/scheduler status from real Deployment state
(list_deployments_for_namespace(), new in modules/kubernetes.py)
instead of `docker inspect`.
- Health probe now hits the real domain directly from the pod (it has
normal internet egress) instead of SSH-ing back out to the host to
curl a public HTTPS URL, which is what the RUNNING_ON_MAIN_SERVER
branch would have done for a pod whose hostname never matches the
main-server hostname check.
Same API/template contract as before (get_sites_list/get_site_health
return the same field shapes) - templates/pages/sites.html needs no
changes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017avLHFqkiti3g62Anq9sVA
New modules/kubernetes.py lists pods and deployments across the 8 app
namespaces via the in-cluster management-platform-viewer-sa (read-only,
no secrets/exec/log), with metrics-server usage as a best-effort extra
that degrades silently when unreachable. New /cluster page and
/api/cluster endpoint, nav entry, and templates render pods/deployments
grouped by namespace using the existing card/badge design system.
Additive only — no existing Docker routes or views touched.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>