Two gaps found testing this live from the standby:
1. restore.html's "Restore on This Server" was checked by default
regardless of RUNNING_ON_MAIN_SERVER — on the standby that option is
nonsensical (no local cluster/kubectl at all) and restore_start()
would have just tried and failed confusingly. Now: that radio is
disabled with an explanatory note when not on the main server,
"External Machine" is checked instead and pre-filled with the tunnel
details (localhost:2224, contabo-key) so restoring from the standby
just targets the real main server without the user having to know
any of that. Added a matching server-side guard in restore_start()
for target=='local' + not RUNNING_ON_MAIN_SERVER (defense in depth —
the UI already prevents it, this catches a direct API call too).
Also fixed refreshSystemMetrics() in platform.js, which would have
overwritten the disabled option's label with the (now-correct, see
previous commit) main-server hostname — looking like a working local
target when it isn't.
2. sync-standby-platform.sh only ever mirrored platform/ — but
restore_start() references /root/CloudOps/backup/restore-k8s-apps.sh
as a fixed absolute path to scp to the remote target, and that
directory never existed on the VM at all. Every restore attempt from
the standby failed immediately with "restore-k8s-apps.sh not found",
regardless of target. Now mirrors /root/CloudOps/backup/ too.
Verified end-to-end for real: triggered a restore of frappe/erpnext from
the standby's actual web UI (target=remote, localhost:2224) — connected
over the tunnel, copied the backup archive + script to the main server,
ran restore-k8s-apps.sh there, scaled the deployment down/up, restored
the DB. Confirmed after: 740 tables in the DB, /api/method/ping
responding on the live pod. Not a dry run — a real restore, actually
initiated from the standby machine.
The VM standby's management-platform was found 2+ months stale during
failover investigation (last deploy: June 10). Root cause: the old
pull-platform-backup.sh + deploy-platform.sh pair on the VM pulled/
redeployed platform-backup-*.tar.gz tarballs of /root/management-platform,
but tarball generation was disabled 2026-08-14 when the source of truth
moved to git-tracked /root/CloudOps — nothing replaced it, so the VM kept
re-deploying its last cached snapshot every hour indefinitely.
New script runs hourly from the main server (cron) instead of pulling from
the VM — pushes /root/CloudOps/platform directly via rsync, plus the live
/root/management-platform/config.py (gitignored, holds real secrets/DB
creds — copied byte-for-byte since it already contains the
RUNNING_ON_MAIN_SERVER auto-detect logic that makes it work correctly
unmodified on both servers). Only restarts the systemd service when rsync
or config actually changed. Old pull/deploy cron entries commented out on
the VM (superseded, not deleted, for history).
Verified end-to-end: ran it live, standby now serves current code (status-
dot markers present in synced templates/app.py), service restarted clean,
config auto-detect correctly logs "running on VM / backup host" and takes
the SSH-fallback code path as designed.