Give management-platform a scoped ServiceAccount for k8s backup/restore

The pod running app.py had no kubectl binary and only a read-only SA
(management-platform-viewer-sa) — backup-k8s-apps.sh/restore-k8s-apps.sh
could never actually run from inside it. Adds:

- management-platform-backup-role (ClusterRole, bound via RoleBinding only
  in n8n/odoo/mautic/nextcloud/erpnext — not cluster-wide): get/list/watch
  on pods/deployments/configmaps/ingresses/PVCs, pods/exec create, pods
  create+delete (needed for restore's populate-before-scale-up loader pod),
  secrets get/list/update/patch, deployments/scale update/patch only (no
  write on full Deployment/Service/ConfigMap/Ingress specs).
- management-platform-sa, replacing management-platform-viewer-sa as the
  pod's identity — inherits the old viewer-role's read-only bindings too
  (retargeted in management-platform-viewer-rbac.yaml) so the existing
  cluster-view page keeps working under one SA.
- k3s binary hostPath-mounted read-only into the pod (same pattern as the
  Jenkins agent's docker-cli container), so `k3s kubectl` is available
  where no separate kubectl binary exists.

Both scripts updated to fall back to k3s kubectl when no kubectl is on
PATH, and to use /proc/meminfo instead of `free` for the resource-safety
check (not present in the pod's minimal image). Also fixes a real issue
found while live-testing this from inside the pod: the pod's pre-existing
/root hostPath mount also exposes the host's own admin ~/.kube/config,
which k3s kubectl was silently preferring over the scoped SA token —
forcing --kubeconfig=/dev/null in both scripts closes that.

Verified end-to-end from inside the actual management-platform pod:
manifests/secret/DB-dump/PVC-data backup for n8n succeeds under the new
SA's scoped permissions, and a kube-system access attempt is correctly
rejected (Forbidden) once the kubeconfig leak is closed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
root
2026-08-20 11:30:03 +02:00
parent 8255527254
commit f3f08c3ef6
5 changed files with 222 additions and 25 deletions

View File

@@ -36,6 +36,26 @@ set -uo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
cd "$SCRIPT_DIR"
# When run inside the management-platform pod, only the static k3s binary is
# hostPath-mounted (no separate kubectl binary) — fall back to `k3s kubectl`.
# IMPORTANT: this pod also has a pre-existing rw hostPath mount of /root (for
# unrelated file-browsing features), which means the host's own root-owned
# ~/.kube/config (full cluster-admin) is silently visible inside the pod too.
# `--kubeconfig=/dev/null` is REQUIRED here to stop k3s kubectl from picking
# that up as a base config — without it, every command below silently runs
# as cluster-admin instead of the scoped management-platform-sa token,
# defeating the RBAC in management-platform-backup-rbac.yaml entirely.
if ! command -v kubectl &>/dev/null && command -v k3s &>/dev/null; then
kubectl() {
k3s kubectl \
--kubeconfig=/dev/null \
--server=https://kubernetes.default.svc \
--certificate-authority=/var/run/secrets/kubernetes.io/serviceaccount/ca.crt \
--token="$(cat /var/run/secrets/kubernetes.io/serviceaccount/token)" \
"$@"
}
fi
APPS_FILTER="all"
VALID_APPS="frappe odoo nextcloud mautic n8n"
@@ -98,7 +118,9 @@ echo "========================================="
check_resources() {
local avail_mb avail_gb
avail_mb=$(free -m | awk '/^Mem:/{print $7}')
# /proc/meminfo instead of `free` — `free` isn't present in the
# management-platform pod's minimal image, /proc/meminfo always is.
avail_mb=$(awk '/^MemAvailable:/{printf "%d", $2/1024}' /proc/meminfo)
avail_gb=$(df -BG / | awk 'NR==2{gsub("G","",$4); print $4}')
if [ "$avail_mb" -lt 1536 ] || [ "$avail_gb" -lt 10 ]; then
echo " ❌ Resource safety threshold hit (RAM=${avail_mb}MB, Disk=${avail_gb}GB) — aborting."
@@ -139,15 +161,21 @@ for app in $ALL_APPS; do
app_ok=true
# ---- 1. Secret + manifests (apply — idempotent) ----
# ---- 1. Secret (apply — idempotent update, no create: every Secret
# restored here already exists in a same-cluster restore) ----
if [ -f "$APP_DIR/secret.yaml" ]; then
echo -n " 🔑 Applying Secret ... "
kubectl apply -f "$APP_DIR/secret.yaml" &>/dev/null && echo "✅" || { echo "⚠️ FAILED"; app_ok=false; }
fi
if [ -f "$APP_DIR/manifests.yaml" ]; then
echo -n " 📄 Applying manifests ... "
kubectl apply -f "$APP_DIR/manifests.yaml" &>/dev/null && echo "✅" || { echo "⚠️ FAILED"; app_ok=false; }
fi
# NOTE: manifests.yaml (Deployment/Service/ConfigMap/Ingress specs) is
# captured by backup-k8s-apps.sh but deliberately NOT re-applied here.
# Same-cluster restore-in-place never needs to reconcile those specs —
# only Secret values, PVC data, and DB content actually change — and
# the pod's RBAC (management-platform-backup-role) intentionally grants
# no write on Deployments/Services/ConfigMaps/Ingresses beyond the
# deployments/scale subresource used in step 2/4 below. manifests.yaml
# is kept in every backup purely as a captured reference for a future
# fresh-cluster/DR restore path, which is explicitly out of scope today.
# ---- 2. Scale app down ----
prior_replicas=$(kubectl get deployment "$app_deploy" -n "$ns" -o jsonpath='{.spec.replicas}' 2>/dev/null || echo 1)