K8s / argo stack for my more uptime-sensitive services
  • HCL 68.1%
  • Makefile 19%
  • Shell 10.9%
  • Dockerfile 2%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-21 15:22:24 -04:00
apps Migrate argo to new git host 2026-09-17 17:19:33 -04:00
charts Update domain for repertory 2026-09-21 15:22:24 -04:00
images/mcp-auth-proxy Migrate ECR to forgejo container registry 2026-08-22 17:56:03 -04:00
manifests Hetzner migration 2026-09-21 13:07:52 -04:00
scripts Hetzner migration 2026-09-21 13:07:52 -04:00
terraform Update domain for repertory 2026-09-21 15:22:24 -04:00
.gitignore Hetzner migration 2026-09-21 13:07:52 -04:00
.goosehints Hetzner migration 2026-09-21 13:07:52 -04:00
.sops.yaml Hetzner migration 2026-09-21 13:07:52 -04:00
AGENTS.md Hetzner migration 2026-09-21 13:07:52 -04:00
controlplane.yaml Terraform all resources, update encryption scheme, add ingress 2026-08-17 10:22:06 -04:00
Makefile Hetzner migration 2026-09-21 13:07:52 -04:00
opencode.jsonc Revert "Switch from athena to grafana / loki" 2026-09-21 10:21:26 -04:00
README.md Hetzner migration 2026-09-21 13:07:52 -04:00
talosconfig Terraform all resources, update encryption scheme, add ingress 2026-08-17 10:22:06 -04:00
values.yaml Update registry host from forge.keane.sh to git.keane.sh 2026-09-20 13:41:46 -04:00
worker.yaml Terraform all resources, update encryption scheme, add ingress 2026-08-17 10:22:06 -04:00

dumpnet-argo

GitOps cluster management for dumpnet — a single-node Talos/Kubernetes cluster on Hetzner Cloud, managed via ArgoCD. (Migrated off AWS EC2; AWS is still used for Route53, SES, and one S3/CloudFront static site.)

Architecture

  • Talos Linux on a Hetzner Cloud server — immutable, API-driven OS
  • ArgoCD — GitOps continuous delivery, self-managing
  • ingress-nginx — ingress controller (hostNetwork mode)
  • App of Apps pattern — all workloads defined in manifests/, grouped under apps/
  • Terraform — cluster infrastructure and bootstrap (Hetzner compute + a handful of AWS services)
  • SOPS + age — application secrets, committed encrypted directly in git, decrypted in-cluster by sops-secrets-operator

Prerequisites

  • terraform
  • kubectl
  • talosctl
  • sops + age (age-keygen)
  • hcloud-upload-image (one-time Talos image upload — see below)
  • AWS CLI configured (aws configure) — only needed for Route53/SES/S3, not compute
  • A Hetzner Cloud API token (TF_VAR_hcloud_token env var — never put this in a tracked file)

Secrets Setup (first time only)

Two independent secrets mechanisms, don't conflate them:

  • Application secrets (Postgres password, API keys, etc.) — SOPS + age, encrypted directly in git as SopsSecret CRs (charts/*/templates/sops-secret.yaml). Decrypted in-cluster by sops-secrets-operator, which needs the age private key as a bootstrap kubectl secret (see step 5 below) — this can't be automated by Terraform/ArgoCD, since nothing that could read the private key should be applied automatically.
  • Cluster credentials (kubeconfig, talosconfig, ArgoCD admin password) — come straight from Terraform outputs (terraform output -raw ...), never stored anywhere else. controlplane.yaml/worker.yaml/talosconfig local files are gitignored entirely, never committed in any form (encrypted or not).

Generate an age key once (if you don't already have one):

mkdir -p ~/.config/sops/age
age-keygen -o ~/.config/sops/age/keys.txt
# public key goes in .sops.yaml (already set for this repo's key);
# back up the private key file securely, it's the only copy

First-Time Setup

1. Clone and init

git clone https://git.keane.sh/ian/dumpnet-argo.git
cd dumpnet-argo
export TF_VAR_hcloud_token=...   # Hetzner Cloud API token
make init

2. Build the Talos image (one-time, or on a Talos version bump)

Hetzner's own docs mention a public Talos ISO they host themselves — but checking a real account's ISO list shows it doesn't actually exist (confirmed the hard way). So build/upload one instead:

export HCLOUD_TOKEN=$TF_VAR_hcloud_token
./scripts/build-talos-hetzner-image.sh
# note the resulting snapshot ID

3. Configure variables

cp terraform/terraform.tfvars.example terraform/terraform.tfvars
# edit terraform/terraform.tfvars: hosted_zone_id, and hetzner_image
# (the snapshot ID from step 2)

4. Spin up the cluster

make apply

Note: make apply runs two Terraform passes internally. The first pass (-target=talos_cluster_kubeconfig.this) brings up the Hetzner server, bootstraps the cluster, and retrieves the kubeconfig. The second pass then uses that live kubeconfig to provision Kubernetes resources (namespaces, ArgoCD Helm release). This two-phase approach is necessary because the Helm and Kubernetes Terraform providers need a reachable cluster to initialize.

This will:

  • Allocate a Hetzner Primary IP (allocated independently of the server, so Talos can know its own endpoint upfront — no EIP-association dance needed)
  • Create a Hetzner Firewall (Talos API, Kubernetes API, HTTP/HTTPS, RTMP, WebRTC)
  • Launch a Hetzner Cloud server from the Talos snapshot, with machine config as user-data
  • Apply the machine configuration via the Talos API
  • Bootstrap etcd
  • Retrieve the kubeconfig
  • Install ArgoCD via Helm with a generated admin password
  • Create Route53 DNS records pointing at the Hetzner Primary IP

5. Bootstrap the SOPS age key (one-time, per cluster)

The sops-secrets-operator can't decrypt anything until it has the age private key:

kubectl create secret generic sops-age-key-file \
  --namespace sops \
  --from-file=key=$HOME/.config/sops/age/keys.txt

6. Post-apply bootstrap

make post-apply

This script:

  • Fetches kubeconfig and talosconfig from Terraform outputs
  • Waits for the node to be ready
  • Applies the ArgoCD App of Apps
  • Builds/pushes the mcp-auth-proxy image

7. Get ArgoCD password

make argocd-password

Then log in at https://argocd.dumpnet.chat — ArgoCD will finish deploying everything else automatically.

Rebuilding from Scratch

To fully tear down and recreate the cluster:

make clean    # clear stale kube/talos contexts
make destroy  # tear down all terraform resources
make apply    # recreate everything
# re-run steps 5-7 above (age key bootstrap, post-apply, get password)

Day-to-Day

  • Add a new app: add a manifest to manifests/ and a chart to charts/ — ArgoCD picks it up on next sync
  • Add a DNS record: add the subdomain to dns_records in terraform/terraform.tfvars and run make apply
  • Add/change a secret: edit the plaintext, sops --encrypt --in-place charts/<name>/templates/sops-secret.yaml, never commit the plaintext version
  • Cluster access: make kubeconfig or make talosconfig (reads straight from Terraform outputs, safe to re-run — automatically clears any stale context from a previous cluster before merging)

Makefile Reference

Command Description
make apply Create/update cluster infrastructure
make plan Preview infrastructure changes
make destroy Tear down everything
make clean Pre-destroy cleanup (clear stale kube/talos contexts)
make post-apply One-time bootstrap after fresh cluster creation
make bootstrap Apply App of Apps only
make kubeconfig Pull kubeconfig from Terraform output (clears stale context first)
make talosconfig Pull talosconfig from Terraform output
make argocd-password Print ArgoCD admin password (from Terraform output)
make build-mcp-auth-proxy Build/push the mcp-auth-proxy image
make build-bl0t Build/push the bl0t runtime image (rarely needed)
make stream-url Print the current RTMP publish URL

Repo Structure

dumpnet-argo/
├── Makefile                        # top-level commands
├── .sops.yaml                      # SOPS age recipient + encryption rules
├── scripts/
│   ├── post-apply.sh                # one-time bootstrap after fresh cluster
│   └── build-talos-hetzner-image.sh # one-time Talos snapshot upload
├── apps/                           # app-of-apps groups (cluster/data/services/mcp)
├── manifests/                      # ArgoCD Application manifests, grouped to match apps/
├── charts/                         # Helm charts + values per app
│   └── sops-secrets-operator/      # decrypts SopsSecret CRs in-cluster
└── terraform/                      # cluster infrastructure
    ├── hetzner.tf                   # server, firewall, primary IP, Talos bootstrap
    ├── versions.tf
    ├── variables.tf
    ├── argocd.tf
    ├── dns.tf
    ├── ses.tf
    ├── repertory-frontend.tf         # S3/CloudFront static site (still AWS)
    └── outputs.tf

Forking / Multiple Environments

Global per-cluster config lives in values.yaml at the repo root:

clusterName: dumpnet
domain: dumpnet.chat
repoURL: https://git.keane.sh/ian/dumpnet-argo.git
certEmail: dumpnetcerts@keane.sh
awsRegion: us-east-1

This file is passed as the first valueFiles entry to every Helm chart, so domain, certEmail, clusterName, and awsRegion are available as {{ .Values.* }} in all chart values. Terraform variables in terraform/variables.tf mirror these same settings for the infrastructure side.

The one thing that can't be templated is repoURL in the ArgoCD Application manifests themselves (under apps/ and manifests/). These are plain YAML consumed by ArgoCD before any Helm rendering happens — there's no way to interpolate them without a Config Management Plugin. When forking this repo, do a global find/replace on git.keane.sh/ian/dumpnet-argo with your own repo URL.

Notes

  • The control plane taint is disabled via allowSchedulingOnControlPlanes: true in the Talos machine config — no manual taint removal needed
  • ArgoCD manages itself after initial Helm install — future upgrades go through the repo, and both the Helm release (terraform/argocd.tf) and ArgoCD's own self-managed Application (manifests/cluster/argocd.yaml) should be pinned to the same chart version, or they'll fight each other (this bit us once)
  • Sensitive local files (controlplane.yaml, worker.yaml, talosconfig) are gitignored entirely — never committed, encrypted or not
  • Logs: fluent-bit currently discards output (null sink) — the previous S3+Athena setup was retired, and the Loki/Grafana replacement was reverted after contributing to an OOM incident on the old 4GB node. Revisit once there's real headroom.