Two separate bugs found while testing the Audit/Details UI against a
real k8s-format backup (myapps-k8s-backup-*):
1. Format-hardcoded checks. audit_backup()'s Internal Structure and
Volume Count checks assumed the legacy Docker-era archive layout
(volumes/*.tar.gz, compose-files/) and would fail every k8s-format
backup (per-app dirs with manifests.yaml/db-dump.sql.gz/pvc-data.tar.gz)
even when it's perfectly healthy - same root cause as the
myapps-backup-*/myapps-k8s-backup-* prefix bug fixed earlier, just in
the audit checks instead of the listing/delete glob. cloud_backup.py's
r2_audit_backup() had the identical unpatched filename regex, and
app.py's /api/backups/details route had its own, which outright
400'd any k8s-format filename before even looking at it.
2. A bigger, separate bug this surfaced: get_local_backups(),
get_vm_backups(), and _resolve_archive_path() all branch on
RUNNING_ON_MAIN_SERVER to decide between direct filesystem access and
SSH-to-self - and from inside the management-platform pod that flag's
hostname check can be wrong for the pod's own ephemeral hostname
depending on how it's evaluated, sending these down the SSH branch
using a topology (separate warm-standby-on-the-VM SSH path) that
doesn't apply to a pod that already has /root hostPath-mounted in.
Fixed by trying direct filesystem access first wherever the archive
would already be locally reachable, falling back to the existing SSH
paths otherwise - this preserves the original warm-standby-on-VM
failover design (a real second deployment of this platform on the VM
host, kept reachable via SSH when the main server is down) completely
unchanged; it only adds the fast local path for the case where the
caller already has direct filesystem access.
Verified against both a real k8s-format backup (odoo, single-app) and a
real legacy-format backup on disk - both audit correctly now, with
format-appropriate checks and labels.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017avLHFqkiti3g62Anq9sVA
modules/sites.py was still a fully Docker-era hardcoded registry:
container names (odoo-clean-odoo-1, frappe-erpnext, nextcloud-app,
mautic-app, n8n-app) that no longer exist post-migration, checked via
`docker inspect` over SSH, plus hardcoded domain/port fields -
nextcloud and mautic even had domain: None, silently falling back to
dead Docker host-port URLs while their real Ingress domains
(next.cloud.nav.ovh, mautics.nav.ovh) sat unused.
Now sources everything live from the cluster:
- Domain + TLS from the actual Ingress object per app (via the new
get_ingress_info() in modules/kubernetes.py), not a static guess.
Odoo/Nextcloud's Ingress lives in `default` (see prior commit for the
RBAC this needed); n8n/mautic/erpnext's lives in their own namespace.
- App/DB/cache/worker/scheduler status from real Deployment state
(list_deployments_for_namespace(), new in modules/kubernetes.py)
instead of `docker inspect`.
- Health probe now hits the real domain directly from the pod (it has
normal internet egress) instead of SSH-ing back out to the host to
curl a public HTTPS URL, which is what the RUNNING_ON_MAIN_SERVER
branch would have done for a pod whose hostname never matches the
main-server hostname check.
Same API/template contract as before (get_sites_list/get_site_health
return the same field shapes) - templates/pages/sites.html needs no
changes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017avLHFqkiti3g62Anq9sVA
api_backup_run and restore_start were still hardcoded to the Docker-era
backup-myapps.sh/restore-myapps.sh, which target Docker volumes that no
longer exist for the 5 apps now running in k3s (n8n, odoo, mautic,
nextcloud, frappe/erpnext) — every backup/restore triggered from the UI
has been silently hollow for these apps since the migration.
- api_backup_run: /root/backup-myapps.sh -> /root/CloudOps/backup/backup-k8s-apps.sh
- restore_start: switched from the old image-relative path (which resolved
to /app/restore-myapps.sh, baked into the Docker image from platform/)
to an absolute /root/CloudOps/backup/restore-k8s-apps.sh path — the new
backup/ folder is a sibling of platform/, not part of the Docker build
context, so it's only reachable via the pod's existing /root hostPath
mount, the same way api_backup_run already reaches its script.
--apps flag shape and the job-log streaming contract are unchanged, so no
frontend/template changes needed for the core flow. modules/backups.py's
three hardcoded myapps-backup-* references (get_local_backups, get_vm_backups
x2, delete_backup's filename regex) now also recognize the new
myapps-k8s-backup-* prefix, so both backup lineages stay visible/manageable
in the UI.
Depends on the previous commit (scoped ServiceAccount + k3s kubectl access
for the pod) to actually function when triggered from the UI.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
New modules/kubernetes.py lists pods and deployments across the 8 app
namespaces via the in-cluster management-platform-viewer-sa (read-only,
no secrets/exec/log), with metrics-server usage as a best-effort extra
that degrades silently when unreachable. New /cluster page and
/api/cluster endpoint, nav entry, and templates render pods/deployments
grouped by namespace using the existing card/badge design system.
Additive only — no existing Docker routes or views touched.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>