Unify restore + DR bootstrap: every remote restore now goes through Ansible
Previously restore_start()'s 'remote' target did its own manual scp+ssh+ restore-k8s-apps.sh dance, entirely separate from the new DR bootstrap playbook (ansible/dr-bootstrap.yml) — meaning restoring onto a genuinely empty server would just fail (no namespace/PVC/Deployments, and restore-k8s-apps.sh assumes those exist). Every "External Machine" restore now runs the same playbook instead: it's safe for both cases, not just the empty-server one — create-namespace-if-missing and `kubectl apply` for the sanitized PVC/Secret/manifests are no-ops against a target that already has this app running with a matching spec (apply only reconciles differences, and a backup's own captured manifests are by definition identical to what's already live) — so "restore onto an existing cluster" and "restore onto nothing" are the same command now; the playbook's own checks-then-acts steps decide how much of it actually needs to do anything. Three real bugs found getting this working, not just wiring it up blind: 1. dr-bootstrap.yml's `hosts: dr_target` only matches a named inventory group — the dynamic single-host inventory app.py builds per-request (`-i '<ip>,'`) doesn't create one, so nothing matched and the play silently skipped. Changed to `hosts: all`, which both invocation styles satisfy. 2. ansible-playbook is a pip console-script installed next to whichever python is running — the main pod's system python (no venv there) or this server's own venv/bin on the standby. A bare "ansible-playbook" in the shelled-out command only resolves on the main pod; the standby's venv/bin is never on PATH when its python is invoked directly rather than through `activate`, so it'd fail there. Resolved relative to sys.executable instead, which is correct in both. 3. Passing connection options via `-e ansible_ssh_common_args='-o StrictHostKeyChecking=no ...'` hit a real bug in this ansible-core version's SSH connection plugin: its own internal tty-detection re-parses that string with a strict argparse and throws "argument -o: expected one argument" even for one well-formed -o KEY=VALUE — verified directly on the CLI, not a shell-quoting artifact from this code. Switched to ANSIBLE_HOST_KEY_CHECKING=False / ANSIBLE_TIMEOUT=15 env vars, ansible's own dedicated mechanism for the same effect, which bypasses that code path entirely. Also added ansible-core to requirements.txt (installed automatically by sync-standby-platform.sh's existing `pip install -r requirements.txt` step; needs a Jenkins rebuild to reach the main pod's image), and synced ansible/ to the standby the same way backup/ already was — restore_start() references dr-bootstrap.yml as a fixed absolute path, and that directory didn't exist there at all before this. Verified for real end-to-end: triggered a restore via the standby's actual web UI (target=remote, localhost:2224 tunnel) — Ansible ran env checks, found k3s already present, reconciled the namespace/PVC/Secret/ manifests (all no-ops against the live cluster), then restore-k8s-apps.sh restored the data. n8n on the real main server came back healthy (healthz ok) with all 11 workflows intact in the restored Postgres DB.
This commit is contained in:
@@ -44,6 +44,13 @@ LIVE_CONFIG="/root/management-platform/config.py"
|
||||
# here to scp across, so mirror it too, not just platform/.
|
||||
LOCAL_BACKUP_SRC="/root/CloudOps/backup/"
|
||||
VM_BACKUP_DIR="/root/CloudOps/backup"
|
||||
|
||||
# Same reason as backup/ above — restore_start() references this as a
|
||||
# fixed absolute path (/root/CloudOps/ansible/dr-bootstrap.yml) to drive
|
||||
# every "External Machine" restore now, not just true from-scratch DR.
|
||||
LOCAL_ANSIBLE_SRC="/root/CloudOps/ansible/"
|
||||
VM_ANSIBLE_DIR="/root/CloudOps/ansible"
|
||||
|
||||
LOG="/root/CloudOps/scripts/sync-standby-platform.log"
|
||||
|
||||
SSH_OPTS="-i $VM_KEY -p $VM_PORT -o StrictHostKeyChecking=no -o ConnectTimeout=15"
|
||||
@@ -82,6 +89,17 @@ if [ $? -ne 0 ]; then
|
||||
fi
|
||||
CHANGES="${CHANGES}${BACKUP_CHANGES}"
|
||||
|
||||
ssh $SSH_OPTS "$VM_USER@$VM_HOST" "mkdir -p $VM_ANSIBLE_DIR"
|
||||
ANSIBLE_CHANGES=$(rsync -a --delete -e "ssh $SSH_OPTS" \
|
||||
--exclude tmp --exclude '*.retry' \
|
||||
--itemize-changes \
|
||||
"$LOCAL_ANSIBLE_SRC" "$VM_USER@$VM_HOST:$VM_ANSIBLE_DIR/" 2>>"$LOG")
|
||||
if [ $? -ne 0 ]; then
|
||||
log "❌ rsync of ansible/ failed"
|
||||
exit 1
|
||||
fi
|
||||
CHANGES="${CHANGES}${ANSIBLE_CHANGES}"
|
||||
|
||||
CONFIG_CHANGED=""
|
||||
if ! scp $SCP_OPTS "$LIVE_CONFIG" "$VM_USER@$VM_HOST:$VM_DEPLOY_DIR/config.py.new" &>>"$LOG"; then
|
||||
log "❌ Could not copy live config.py to VM"
|
||||
|
||||
Reference in New Issue
Block a user