Files
infra/tofu/svc-monitoring.tf
DmitryandClaude Sonnet 5 a3fe8031fe feat: migrate all managed LXC provisioning to OpenTofu (blue-green)
Blue-green: a new container is created beside the old one, data is copied, the
IP is moved onto it, and the old container is kept stopped as rollback for at
least a week. Keeping the IP means only the VMID changes, and its consumers
(backup jobs, backup audit) already derive it from the registry.

Batch 1 (2026-09-02): emergency-bot 148->151, docker-test 145->152,
gitea 141->153, vaultwarden 140->154, monitoring 146->155, gyro 150->156,
grimmory 149->157.
Batch 2 (2026-09-03): adguard 144->158, mihomo 143->159, ovpn-mini 132->160.
All migratable LXC are now provisioner: tofu. hermes-ai (frozen) and pbs stay.

- tofu/services.tf + tofu/svc-*.tf: one resource per service, reproducing the
  pct-config etalon. /dev/fuse -> features.fuse; /dev/net/tun ->
  device_passthrough (first live use on mihomo and ovpn-mini); gitea bind mount
  -> datastore volume (data finally reaches PBS); console { type = "shell" }
  declared explicitly (provider tracks cmode there).
- services.yml: vmid + provisioner: tofu for every migrated service; features
  strings and device notes updated to the tofu representation; also drops the
  memoir-bot entry and adds homelab_reverse_proxy_image/_unit.
- pve-*.yml: configuration play target is `{{ pve_config_target | default(...) }}`
  so it can run against <name>-new on a temp address (a bare --limit zeroes the
  play instead of retargeting it). Container-creation plays are gated behind
  `provisioner != 'tofu'` / `pve_provisioning_enabled` (meta: end_play), so a
  stray run cannot pct start a stopped OLD VMID on a live IP. Override for
  intentional legacy rollback: -e pve_<svc>_legacy_provisioning_enabled=true.
- ssh_config: drop memoir-bot; ovpn-mini gets ProxyJump none (a jump via ru-vps
  would route through the very tunnel ovpn-mini terminates).
- gyro.yml / uptime-kuma.yml: same pve_config_target override.
- roles/uptime_kuma: only freeze homelab-monitoring when the unit actually
  exists (a fresh blue-green container never had it).
- offsite-restic-yadisk.yml: the gitea restic profile now runs inside the LXC
  (hosts: gitea), since the bind-mount host path is gone after the volume move;
  lost+found excluded (unreadable in an unprivileged LXC, restic exit 3).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012uoq5AVK8mkBgg83Mq6o5V
2026-09-03 07:05:46 +03:00

108 lines
4.2 KiB
Terraform
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ============================================================================
# Сервис №5: monitoring (OLD vmid 146, cloud-pc).
# Эталон снят 2026-09-02: ssh cloud-pc sudo pct config 146 (совпадает с
# cat /etc/pve/lxc/146.conf дословно).
#
# vm_id 155 — новый VMID для blue-green переезда, TMPIP 192.168.1.14
# используется только на шаге 4.3 (проверка перед cutover). После шага 4.5.15
# адрес — боевой 192.168.1.30 (тот же, что и у OLD: в этой схеме переезжает
# VMID, а не IP), OLD (146) остановлен и остаётся откатом минимум неделю.
#
# Активный сервис — Uptime Kuma: SQLite в /opt/uptime-kuma/data/kuma.db,
# UI-состояние (мониторы, уведомления) в Ansible не описано, каталог
# переносится целиком на шаге 4.4. Стек Prometheus + Alertmanager + Grafana
# в этом контейнере ЗАМОРОЖЕН (docs/ai/legacy-warning.md) — его данные лежат
# отдельно в /opt/monitoring (в основном TSDB Prometheus) и в новый контейнер
# намеренно НЕ переносятся; решение о судьбе замороженного стека отдельное,
# не часть этого переезда.
#
# ОСОБЕННОСТЬ features: реестр (services.yml) отмечает, что роль создаёт
# контейнер с nesting=1, а keyctl=1 добавляет отдельной задачей `pct set 146
# --features nesting=1,keyctl=1` уже после создания (playbooks/pve-monitoring.yml).
# Здесь воспроизведён ЭТАЛОН как он есть на живом контейнере — оба флага сразу,
# декларативно, без промежуточного шага: в привилегированном режиме root@pam
# (см. tofu/README.md, «Решение: root@pam по паролю») дошаг `pct set` для этого
# и остальных сервисов больше не нужен, features применяются одной фазой.
# ============================================================================
resource "proxmox_virtual_environment_container" "monitoring" {
node_name = "cloud-pc"
vm_id = 155
unprivileged = true
start_on_boot = true
started = true
tags = ["tofu"]
initialization {
hostname = "monitoring"
ip_config {
ipv4 {
address = "192.168.1.30/24"
gateway = "192.168.1.1"
}
}
dns {
servers = ["1.1.1.1"]
}
user_account {
keys = [trimspace(file(pathexpand("~/.ssh/id_ed25519_homelab.pub")))]
}
}
operating_system {
template_file_id = "local:vztmpl/debian-13-standard_13.1-2_amd64.tar.zst"
type = "debian"
}
cpu {
cores = 2
}
memory {
dedicated = 4096
swap = 512
}
disk {
datastore_id = "data"
size = 24
}
# Эталон: features: nesting=1,keyctl=1 — см. ОСОБЕННОСТЬ в комментарии выше.
# Ни fuse, ни device passthrough эталону не нужны (dev0 в pct config нет).
features {
nesting = true
keyctl = true
}
# roles/pve_lxc всем контейнерам ставит cmode=shell (pve_lxc_cmode), провайдер
# это отслеживает через блок console.type — без него на следующем apply
# откатит на дефолт Proxmox tty. Обнаружено 2026-09-02 на emergency-bot,
# см. tofu/README.md.
console {
type = "shell"
}
network_interface {
name = "eth0"
bridge = "vmbr0"
firewall = true
}
startup {
order = 80
}
}
output "monitoring" {
description = "Что проверять на узле после apply"
value = {
vmid = proxmox_virtual_environment_container.monitoring.vm_id
node = proxmox_virtual_environment_container.monitoring.node_name
verify = "ssh cloud-pc sudo pct config ${proxmox_virtual_environment_container.monitoring.vm_id}"
}
}