feat: migrate all managed LXC provisioning to OpenTofu (blue-green)
Blue-green: a new container is created beside the old one, data is copied, the
IP is moved onto it, and the old container is kept stopped as rollback for at
least a week. Keeping the IP means only the VMID changes, and its consumers
(backup jobs, backup audit) already derive it from the registry.
Batch 1 (2026-09-02): emergency-bot 148->151, docker-test 145->152,
gitea 141->153, vaultwarden 140->154, monitoring 146->155, gyro 150->156,
grimmory 149->157.
Batch 2 (2026-09-03): adguard 144->158, mihomo 143->159, ovpn-mini 132->160.
All migratable LXC are now provisioner: tofu. hermes-ai (frozen) and pbs stay.
- tofu/services.tf + tofu/svc-*.tf: one resource per service, reproducing the
pct-config etalon. /dev/fuse -> features.fuse; /dev/net/tun ->
device_passthrough (first live use on mihomo and ovpn-mini); gitea bind mount
-> datastore volume (data finally reaches PBS); console { type = "shell" }
declared explicitly (provider tracks cmode there).
- services.yml: vmid + provisioner: tofu for every migrated service; features
strings and device notes updated to the tofu representation; also drops the
memoir-bot entry and adds homelab_reverse_proxy_image/_unit.
- pve-*.yml: configuration play target is `{{ pve_config_target | default(...) }}`
so it can run against <name>-new on a temp address (a bare --limit zeroes the
play instead of retargeting it). Container-creation plays are gated behind
`provisioner != 'tofu'` / `pve_provisioning_enabled` (meta: end_play), so a
stray run cannot pct start a stopped OLD VMID on a live IP. Override for
intentional legacy rollback: -e pve_<svc>_legacy_provisioning_enabled=true.
- ssh_config: drop memoir-bot; ovpn-mini gets ProxyJump none (a jump via ru-vps
would route through the very tunnel ovpn-mini terminates).
- gyro.yml / uptime-kuma.yml: same pve_config_target override.
- roles/uptime_kuma: only freeze homelab-monitoring when the unit actually
exists (a fresh blue-green container never had it).
- offsite-restic-yadisk.yml: the gitea restic profile now runs inside the LXC
(hosts: gitea), since the bind-mount host path is gone after the volume move;
lost+found excluded (unreadable in an unprivileged LXC, restic exit 3).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012uoq5AVK8mkBgg83Mq6o5V
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
e349f19e68
commit
a3fe8031fe
@@ -137,11 +137,21 @@
|
||||
delay: 5
|
||||
until: uptime_kuma_health.status == 200
|
||||
|
||||
# На хосте, где замороженный стек никогда не разворачивался (например, новый
|
||||
# контейнер blue-green переезда), юнита homelab-monitoring нет, и systemd-модуль
|
||||
# упал бы с "Could not find the requested service". Смысл шага — "legacy-стек не
|
||||
# должен работать", а на чистом хосте это уже так. Проверено 2026-09-02.
|
||||
- name: Detect the legacy monitoring unit
|
||||
ansible.builtin.stat:
|
||||
path: /etc/systemd/system/homelab-monitoring.service
|
||||
register: uptime_kuma_legacy_unit
|
||||
|
||||
- name: Freeze the legacy Prometheus monitoring stack after Uptime Kuma is healthy
|
||||
ansible.builtin.systemd:
|
||||
name: homelab-monitoring
|
||||
enabled: false
|
||||
state: stopped
|
||||
when: uptime_kuma_legacy_unit.stat.exists
|
||||
|
||||
- name: Remove legacy monitoring firewall rules
|
||||
community.general.ufw:
|
||||
|
||||
Reference in New Issue
Block a user