Files
CloudOps/k8s/monitoring
root a307607a99 Migrate Grafana/Prometheus/SonarQube from docker-compose to k8s (step #10)
Real data migrated, not a fresh start: 1.2GB Prometheus TSDB, Grafana's
existing dashboards/DB (its admin password was already changed from
default — confirmed real, in-use data), SonarQube's data/logs/extensions
and its own Postgres DB (confirmed: the actual `management-platform`
SonarQube project survived, matching the Jenkinsfile's own
-Dsonar.projectKey). Old docker-compose services stopped, not removed —
left as a rollback path.

Real problems hit and fixed along the way (see k8s/monitoring/README.md
for the full detail, worth reading before touching any of this again):

- k3s's bundled Traefik defaults to a LoadBalancer Service, which
  immediately hijacked ports 80/443 via iptables DNAT (klipper-lb) out
  from under this VM's existing unrelated nginx-served app ("nqks") the
  moment k3s was installed. Patched Traefik to ClusterIP via a
  HelmChartConfig; nginx now reverse-proxies the 3 new domains to
  Traefik's ClusterIP instead of Traefik touching host ports at all.

- Caught my own mistake before it caused damage: "corrected" what I
  assumed was a typo (grafanna.nav.ovh -> grafana.nav.ovh) — the two-N
  version is what the user actually specified, and is what points at this
  VM; plain grafana.nav.ovh is a different, unrelated, pre-existing live
  server at 109.199.127.74. Renamed everything back before requesting any
  cert or touching that domain.

- Adding hostPort to the SonarQube pod (for Jenkins' hardcoded
  178.18.243.51:9000) broke pod-to-pod routing to that pod entirely —
  Traefik couldn't reach it (hung/504) while direct host-to-pod and
  host-to-ClusterIP both worked fine the whole time, which is what gave
  it away. Replaced with a host-level socat proxy (systemd unit) to the
  Service's ClusterIP instead, which touches nothing pod-level. Verified
  Jenkins' exact address still works AND Traefik routing to the same pod
  works simultaneously.

- SonarQube needed a DB migration after the version jump
  (DB_MIGRATION_NEEDED), then OOM'd during post-migration rule
  re-registration on the compose file's original 256MB web heap — bumped
  SONAR_WEB_JAVAADDITIONALOPTS.

TLS: real Let's Encrypt certs via certbot (HTTP-01, nginx already owns
port 80 so no DNS-01/OVH-webhook replication needed on this second
cluster), auto-renewing. Access: HTTP Basic Auth in front of all three
domains for browser access; SonarQube also stays reachable unauthenticated
at 178.18.243.51:9000 specifically for Jenkins (which authenticates via
its own token, not BasicAuth — that path would break the scanner),
restricted to the main server's IP via ufw, not open to the internet.

Verified end-to-end: all three https://*.nav.ovh domains return 401
without credentials and load correctly with them; Jenkins' exact
SonarQube address (no auth) still returns 200; real data confirmed
present in Grafana, Prometheus, and SonarQube post-migration.
2026-08-21 16:01:39 +02:00
..

Monitoring stack on k8s (VM: 178.18.243.51)

Migrated from docker-compose (/root/monitoring, /root/CI-CD/sonarqube on the VM — left docker compose stop'd, not removed, as a rollback path) into the VM's own k3s cluster, with real data migrated (not a fresh start): Prometheus TSDB, Grafana dashboards/DB, SonarQube data+DB (confirmed intact: the real management-platform SonarQube project and Grafana's existing admin password both survived).

Apply order

kubectl apply -f monitoring-pvcs.yaml
# migrate data into the PVCs via a loader pod per volume (see chat history —
# not scripted, was a one-time manual migration)
kubectl apply -f monitoring-apps.yaml
kubectl apply -f monitoring-ingress.yaml

Domains — READ THIS BEFORE TOUCHING DNS

  • grafanna.nav.ovh (note: two Nsgrafana.nav.ovh without the typo is a DIFFERENT, unrelated, pre-existing server at 109.199.127.74. Do not point anything at plain grafana.nav.ovh or touch that record.)
  • sonar.nav.ovh
  • prom.nav.ovh

All three already existed as A records pointing at 178.18.243.51 before this work; only nginx-vhost.conf + a Let's Encrypt cert (via certbot, HTTP-01, auto-renews) make them actually resolve to something.

Why nginx is in front of Traefik, not Traefik directly on 80/443

This VM already runs an unrelated app ("nqks") on nginx, owning ports 80/443. k3s's bundled Traefik defaults to a LoadBalancer Service, which on a single-node k3s claims host 80/443 via iptables DNAT (klipper-lb) — this silently hijacked nqks's traffic the first time (found and fixed during this session). Traefik's Service is patched to ClusterIP-only via a HelmChartConfig (kubectl get helmchartconfig -n kube-system traefik); nginx reverse-proxies the 3 domains to Traefik's ClusterIP instead, and handles TLS + BasicAuth itself.

Access control

  • Browser access to all 3 domains: HTTP Basic Auth (/etc/nginx/.htpasswd-monitoring on the VM, user admin) — ask whoever ran this migration for the password, it's not in git.
  • SonarQube is also reachable directly at 178.18.243.51:9000, no auth — this is deliberate, not an oversight. Jenkins' SonarQube server config (main cluster, hudson.plugins.sonar.SonarGlobalConfiguration.xml) is hardcoded to that exact address and authenticates via its own token, not BasicAuth (which would break the scanner). sonarqube-jenkins-proxy.service (systemd, socat) forwards host:9000 -> the sonarqube Service's ClusterIP; ufw restricts port 9000 to the main server's IP only, not the open internet. Do not add BasicAuth in front of this specific path.
    • Fragility note: the socat unit hardcodes the Service's ClusterIP. If the sonarqube Service in the monitoring namespace is ever deleted and recreated (not just the pod — pod recreates keep the same ClusterIP), update the TCP:<ip>:9000 target in /etc/systemd/system/sonarqube-jenkins-proxy.service on the VM to match.

Gotchas hit building this (context for next time)

  • Don't add a k8s Service of type LoadBalancer on this cluster without first patching it to ClusterIP/setting a HelmChartConfig — same hijack risk as Traefik's default.
  • Don't add hostPort to a pod that also needs normal in-cluster (Service/ Ingress) traffic to reach it — it breaks pod-to-pod routing to that pod (confirmed directly: Traefik -> sonarqube hung/504'd the whole time hostPort was present, while direct host-to-pod and host-to-ClusterIP both worked fine — the asymmetry was the tell). Use a host-level proxy (socat) or a NodePort in the 30000-32767 range instead.
  • SonarQube's DB was mid-upgrade (DB_MIGRATION_NEEDED) after the version jump — needed one POST /api/system/migrate_db (admin/admin was already changed on the real instance; used the actual admin session) — and the post-migration rule re-registration OOM'd on the compose file's original 256MB web heap; bumped SONAR_WEB_JAVAADDITIONALOPTS to -Xms512m -Xmx1536m.