Real data migrated, not a fresh start: 1.2GB Prometheus TSDB, Grafana's
existing dashboards/DB (its admin password was already changed from
default — confirmed real, in-use data), SonarQube's data/logs/extensions
and its own Postgres DB (confirmed: the actual `management-platform`
SonarQube project survived, matching the Jenkinsfile's own
-Dsonar.projectKey). Old docker-compose services stopped, not removed —
left as a rollback path.
Real problems hit and fixed along the way (see k8s/monitoring/README.md
for the full detail, worth reading before touching any of this again):
- k3s's bundled Traefik defaults to a LoadBalancer Service, which
immediately hijacked ports 80/443 via iptables DNAT (klipper-lb) out
from under this VM's existing unrelated nginx-served app ("nqks") the
moment k3s was installed. Patched Traefik to ClusterIP via a
HelmChartConfig; nginx now reverse-proxies the 3 new domains to
Traefik's ClusterIP instead of Traefik touching host ports at all.
- Caught my own mistake before it caused damage: "corrected" what I
assumed was a typo (grafanna.nav.ovh -> grafana.nav.ovh) — the two-N
version is what the user actually specified, and is what points at this
VM; plain grafana.nav.ovh is a different, unrelated, pre-existing live
server at 109.199.127.74. Renamed everything back before requesting any
cert or touching that domain.
- Adding hostPort to the SonarQube pod (for Jenkins' hardcoded
178.18.243.51:9000) broke pod-to-pod routing to that pod entirely —
Traefik couldn't reach it (hung/504) while direct host-to-pod and
host-to-ClusterIP both worked fine the whole time, which is what gave
it away. Replaced with a host-level socat proxy (systemd unit) to the
Service's ClusterIP instead, which touches nothing pod-level. Verified
Jenkins' exact address still works AND Traefik routing to the same pod
works simultaneously.
- SonarQube needed a DB migration after the version jump
(DB_MIGRATION_NEEDED), then OOM'd during post-migration rule
re-registration on the compose file's original 256MB web heap — bumped
SONAR_WEB_JAVAADDITIONALOPTS.
TLS: real Let's Encrypt certs via certbot (HTTP-01, nginx already owns
port 80 so no DNS-01/OVH-webhook replication needed on this second
cluster), auto-renewing. Access: HTTP Basic Auth in front of all three
domains for browser access; SonarQube also stays reachable unauthenticated
at 178.18.243.51:9000 specifically for Jenkins (which authenticates via
its own token, not BasicAuth — that path would break the scanner),
restricted to the main server's IP via ufw, not open to the internet.
Verified end-to-end: all three https://*.nav.ovh domains return 401
without credentials and load correctly with them; Jenkins' exact
SonarQube address (no auth) still returns 200; real data confirmed
present in Grafana, Prometheus, and SonarQube post-migration.
73 lines
3.9 KiB
Markdown
73 lines
3.9 KiB
Markdown
# Monitoring stack on k8s (VM: 178.18.243.51)
|
|
|
|
Migrated from docker-compose (`/root/monitoring`, `/root/CI-CD/sonarqube` on
|
|
the VM — left `docker compose stop`'d, not removed, as a rollback path) into
|
|
the VM's own k3s cluster, with real data migrated (not a fresh start):
|
|
Prometheus TSDB, Grafana dashboards/DB, SonarQube data+DB (confirmed intact:
|
|
the real `management-platform` SonarQube project and Grafana's existing
|
|
admin password both survived).
|
|
|
|
## Apply order
|
|
```
|
|
kubectl apply -f monitoring-pvcs.yaml
|
|
# migrate data into the PVCs via a loader pod per volume (see chat history —
|
|
# not scripted, was a one-time manual migration)
|
|
kubectl apply -f monitoring-apps.yaml
|
|
kubectl apply -f monitoring-ingress.yaml
|
|
```
|
|
|
|
## Domains — READ THIS BEFORE TOUCHING DNS
|
|
- `grafanna.nav.ovh` (note: **two Ns** — `grafana.nav.ovh` without the typo
|
|
is a DIFFERENT, unrelated, pre-existing server at 109.199.127.74. Do not
|
|
point anything at plain `grafana.nav.ovh` or touch that record.)
|
|
- `sonar.nav.ovh`
|
|
- `prom.nav.ovh`
|
|
|
|
All three already existed as A records pointing at 178.18.243.51 before this
|
|
work; only `nginx-vhost.conf` + a Let's Encrypt cert (via certbot, HTTP-01,
|
|
auto-renews) make them actually resolve to something.
|
|
|
|
## Why nginx is in front of Traefik, not Traefik directly on 80/443
|
|
This VM already runs an unrelated app ("nqks") on nginx, owning ports 80/443.
|
|
k3s's bundled Traefik defaults to a LoadBalancer Service, which on a
|
|
single-node k3s claims host 80/443 via iptables DNAT (klipper-lb) —
|
|
this silently hijacked nqks's traffic the first time (found and fixed
|
|
during this session). Traefik's Service is patched to ClusterIP-only via
|
|
a HelmChartConfig (`kubectl get helmchartconfig -n kube-system traefik`);
|
|
nginx reverse-proxies the 3 domains to Traefik's ClusterIP instead, and
|
|
handles TLS + BasicAuth itself.
|
|
|
|
## Access control
|
|
- Browser access to all 3 domains: HTTP Basic Auth (`/etc/nginx/.htpasswd-monitoring`
|
|
on the VM, user `admin`) — ask whoever ran this migration for the password,
|
|
it's not in git.
|
|
- **SonarQube is also reachable directly at `178.18.243.51:9000`, no auth** —
|
|
this is deliberate, not an oversight. Jenkins' SonarQube server config
|
|
(main cluster, `hudson.plugins.sonar.SonarGlobalConfiguration.xml`) is
|
|
hardcoded to that exact address and authenticates via its own token, not
|
|
BasicAuth (which would break the scanner). `sonarqube-jenkins-proxy.service`
|
|
(systemd, socat) forwards host:9000 -> the sonarqube Service's ClusterIP;
|
|
ufw restricts port 9000 to the main server's IP only, not the open
|
|
internet. Do not add BasicAuth in front of this specific path.
|
|
- Fragility note: the socat unit hardcodes the Service's ClusterIP. If the
|
|
`sonarqube` Service in the `monitoring` namespace is ever deleted and
|
|
recreated (not just the pod — pod recreates keep the same ClusterIP),
|
|
update the `TCP:<ip>:9000` target in
|
|
`/etc/systemd/system/sonarqube-jenkins-proxy.service` on the VM to match.
|
|
|
|
## Gotchas hit building this (context for next time)
|
|
- Don't add a k8s Service of type LoadBalancer on this cluster without first
|
|
patching it to ClusterIP/setting a HelmChartConfig — same hijack risk as
|
|
Traefik's default.
|
|
- Don't add `hostPort` to a pod that also needs normal in-cluster (Service/
|
|
Ingress) traffic to reach it — it breaks pod-to-pod routing to that pod
|
|
(confirmed directly: Traefik -> sonarqube hung/504'd the whole time
|
|
hostPort was present, while direct host-to-pod and host-to-ClusterIP both
|
|
worked fine — the asymmetry was the tell). Use a host-level proxy (socat)
|
|
or a NodePort in the 30000-32767 range instead.
|
|
- SonarQube's DB was mid-upgrade (`DB_MIGRATION_NEEDED`) after the version
|
|
jump — needed one `POST /api/system/migrate_db` (admin/admin was already
|
|
changed on the real instance; used the actual admin session) — and the
|
|
post-migration rule re-registration OOM'd on the compose file's original
|
|
256MB web heap; bumped `SONAR_WEB_JAVAADDITIONALOPTS` to `-Xms512m -Xmx1536m`.
|