Octov0.11.7
Deployment

GCP with Terraform

The reference GCP deployment: k3s, Traefik, TLS, and per-integration subdomains.

The reference production deployment runs Octo on a single GCP VM with single-node k3s, Traefik ingress, and cert-manager for free auto-renewing Let's Encrypt TLS. The editor is served at your domain and each deployed integration can get its own subdomain. Everything is codified in deploy/terraform/.

Architecture

                      Cloud DNS
   octo.example.com          ─┐
   *.octo.example.com        ─┤ A records → static IP
                              │
                       ┌──────▼─────────────────────────────────────────┐
                       │ GCE VM (Debian 12, e2-standard-2), single-node  │
                       │ k3s                                             │
                       │                                                 │
                       │  Traefik :80/:443  ── cert-manager (Let's       │
                       │    │                   Encrypt, HTTP-01)        │
                       │    ├─ octo.…           → editor  :3000          │
                       │    └─ {slug}.octo.…    → integration pod :8080  │
                       │                                                 │
                       │  editor ─(BFF)→ orchestrator :8090              │
                       │  Postgres :5432 (StatefulSet)                   │
                       │  per integration: ConfigMap + Deployment +      │
                       │    Service [+ Ingress when exposed externally]  │
                       └─────────────────────────────────────────────────┘
                              ▲ images + OCI chart (pull)
                       Artifact Registry  ◄── Cloud Build (on version tag)

Ownership is split three ways:

The VM only bootstraps the cluster: k3s, cert-manager, the letsencrypt-prod ClusterIssuer, and an octo-pull helper. Terraform owns the Octo release through the Helm provider, so you apply the release root to install or upgrade the chart. The orchestrator creates per-integration Kubernetes resources at runtime when you deploy an integration from the editor.

The two Terraform roots

Under deploy/terraform/:

RootWhat it createsWhen to apply
infra/Artifact Registry repo, VM + static IP + firewall + DNS records + k3s bootstrap, and (optional) the Cloud Build trigger + IAMOnce per cluster
release/The Octo Helm release (image tag + chart version)Every deploy or upgrade

Each root owns its own terraform.tfvars (gitignored; copy the committed .example). Terraform loads that filename automatically, so no -var-file is passed anywhere. Per-deploy values (image_tag, chart_version) come from the command line. Both roots keep state in the same versioned GCS bucket, one prefix each, so Cloud Build and your laptop share it.

The release state holds generated secrets (the Postgres password and, with the OIDC client secret and the Auth.js session secret), and the fetched kubeconfig holds cluster-admin credentials. Both are gitignored, so keep the state bucket locked down.

Prerequisites

  • gcloud authenticated: gcloud auth application-default login.
  • An existing Cloud DNS managed zone for your domain (you supply its name, e.g. example-com). Terraform creates the A record for the editor host and the *.{domain} wildcard record pointing at the static IP.
  • On your workstation: terraform (>= 1.5), helm, docker, gcloud, and task.
  • Default region is us-west1; override per root if needed.

One-time setup

Fill in the tfvars file

gcloud auth application-default login
cd deploy/terraform
cp backend.hcl.example backend.hcl        # the shared state bucket
cp infra/terraform.tfvars.example infra/terraform.tfvars
cp release/terraform.tfvars.example release/terraform.tfvars

Set at least project_id, domain, and dns_managed_zone. project_id and domain appear in both files and must agree.

Create the state bucket

Backs the release root's remote state; run once:

task state:bucket PROJECT=<your-project>

Apply the infra root

Creates the registry, VM, static IP, firewall, DNS records, and the k3s bootstrap:

task infra:apply
terraform -chdir=deploy/terraform/infra output   # image_base, static_ip, url, kube_api_endpoint

(Optional) Enable Cloud Build automation

Connect the GitHub repo once in the console (Cloud Build → Triggers → Connect repository, which installs the GitHub App; Terraform cannot do this), then set enable_cloudbuild = true in infra/terraform.tfvars and re-run task infra:apply.

SSH (22) and the k3s API (6443) default to open. In production, set ssh_source_ranges and kube_api_source_ranges to your IP, but include the IAP range 35.235.240.0/20 so the Cloud Build deploy step can still reach the VM. The release apply needs 6443 reachable from wherever Terraform runs.

Required and notable variables

Set in each root's terraform.tfvars (values needed by both must agree):

VariableDefaultNotes
project_id(none)Required.
domainocto.juancavallotti.comEditor host; integrations live under *.{domain}.
dns_managed_zone(none)Required; the Cloud DNS zone name.
machine_typee2-standard-28 GB RAM; e2-medium (4 GB) is tight.
ssh_source_ranges["0.0.0.0/0"]Restrict in production (keep the IAP range).
kube_api_source_rangesnull (= SSH ranges)Who can reach 6443.
acme_email(none)Let's Encrypt account email.
enable_cloudbuildfalseCreate the trigger (GitHub App must be connected first).
cloudbuild_auto_deploytrueAlso roll the cluster on a version tag (_DEPLOY=true).
oidc_client_id / oidc_client_secret / oidc_issuerunsetRequired. The identity provider the editor signs people in against; consumed by the release root. There is no unauthenticated mode, so an apply without them fails at the chart.

The release root additionally takes image_tag (default latest) and chart_version (required, must match the published helm/Chart.yaml), both supplied on the command line by task deploy or Cloud Build.

Where the credentials live

There is no Secret Manager in this setup. The release root generates the Postgres password, the Auth.js session secret, the KV encryption key and the dev-run HMAC key, and holds them in its own state: the versioned, encrypted GCS bucket from backend.hcl. The OIDC client secret and the embedding provider key are supplied by you and persisted next to that state, so a Cloud Build deploy (which has no terraform.tfvars) can read them back.

None of them are handed to Helm. modules/octo-secrets installs each one into the cluster as a Kubernetes Secret before the chart is applied, and modules/helm-release passes only the Secret's name and key:

SecretChart value it satisfies
octo-db-passwordpostgres.auth.existingSecret
octo-auth-credsauth.existingSecret
octo-kv-keykv.existingSecret
octo-devrun-keyorchestrator.devRuns.existingSecret
octo-embeddings-keyembeddings.existingSecret

Helm keeps every value it is given in the release history. set_sensitive marks a value secret to Terraform's plan output and does nothing about the release Secret Helm writes it into, which is retained per revision, in the cluster, and outlives the rotation that was supposed to end it. A reference is the only form that does not leave a copy behind.

Upgrading an installation created before this change. The namespace used to be created by Helm and is now a Terraform resource, so the first apply will try to create one that already exists. Import it once:

terraform -chdir=deploy/terraform/release import \
  'module.secrets.kubernetes_namespace_v1.octo[0]' octo

The credentials come from the same state, so the apply replaces chart values with references and rolls the Deployments once. Revisions written before the change still hold what they were given; helm history and a truncated --history-max retire them.

Publish images and the charts

The node and the Helm provider pull from Artifact Registry, so images and charts must be published before a deploy.

Automated (recommended): push a version tag. release-please publishes vX.Y.Z, and the Cloud Build trigger runs cloudbuild.yaml, building all eleven images and both charts. The images are pushed twice, under the git tag and latest; the charts are pushed only under their own version, because an OCI chart's tag is its version. See Releases and upgrades.

Manual, with IMAGE_BASE from terraform output image_base:

gcloud auth configure-docker us-west1-docker.pkg.dev
helm registry login us-west1-docker.pkg.dev -u oauth2accesstoken -p "$(gcloud auth print-access-token)"
task images:push IMAGE_BASE=$IMAGE_BASE TAG=v0.1.1
task helm:push   IMAGE_BASE=$IMAGE_BASE

helm:push publishes both the octo application chart and the octo-common library chart it is built from.

Deploy and roll upgrades

On a version tag, Cloud Build deploys automatically (the deploy step in cloudbuild.yaml, gated on _DEPLOY). To deploy an already-published tag by hand:

task deploy TAG=v0.1.1            # optional: DOMAIN=… INSTANCE=… ZONE=…

This fetches a fresh kubeconfig from the VM (task deploy:kubeconfig), derives chart_version from helm/Chart.yaml, and applies the release root. Each apply:

  1. Runs octo-pull on the node over SSH (with a fresh registry token) to pull the target image tag into containerd, so the chart's pods (imagePullPolicy: IfNotPresent), including the per-integration runtime image, find it locally.
  2. Installs or upgrades the chart through the Helm provider, passing the Postgres password held in state and authenticating the OCI chart pull with your GCP token.

A changed image_tag rewrites the pod templates, so the Deployments roll automatically; Postgres is untouched when only the tag moves. The first TLS issuance takes a minute after the DNS A record resolves.

Images are pinned by digest on the Cloud Build path

Cloud Build reads back the digest of every image it pushed and renders dist/values.images.yaml pinning them, then passes it to the release root as image_values_file. The release then runs exactly what that build produced, rather than whatever the tag resolves to at apply time, and image_tag is not passed to the chart at all when digests are supplied, so there is only ever one source of truth for what is running. The file is archived to gs://octo-tfstate-PROJECT/release/values.images-TAG.yaml beside the release.

A manual task deploy TAG=… leaves image_values_file empty and identifies images by tag, which works because it also pre-pulls that tag onto the node. To pin a manual deploy, render the file yourself:

task helm:values:images IMAGE_BASE=$IMAGE_BASE TAG=v0.1.1
terraform -chdir=deploy/terraform/release apply \
  -var image_values_file=$PWD/dist/values.images.yaml ...

Database durability

The VM's Postgres volume is provisioned by k3s's local-path provisioner, whose per-claim directory is named after the claim's UID. Re-bootstrapping the VM mints a new UID, bringing the database back empty with the old data stranded on disk. Set postgres_host_path (e.g. /var/lib/octo-data) to pin it to a fixed path.

On a release that already holds data this is a data move, not a config change: copy the directory across with the StatefulSet scaled to zero first, or the database comes up empty. See docs/deployment.md for the procedure. Either way the data sits on the boot disk and does not survive the VM being destroyed; an attached data disk mounted at the pinned path does.

Integration endpoints

When you deploy an integration from the editor, the orchestrator creates its workload in octo-dev, reachable internally at http://octo-int-{slug}.octo-dev:8080. Toggling Expose externally adds a Traefik Ingress at https://{subdomain}.{domain}, with TLS issued per host by cert-manager via HTTP-01 (the wildcard DNS record resolves every subdomain to the VM). External endpoints require BASE_DOMAIN on the orchestrator, which the release root sets to your domain. See Deployments for the full endpoint model.

Operations

# SSH to the VM
gcloud compute ssh octo --zone us-west1-a

# Cluster state
gcloud compute ssh octo --zone us-west1-a -- sudo k3s kubectl get pods -n octo-dev

# Platform / orchestrator logs (on the VM)
sudo k3s kubectl logs -n octo-dev deploy/octo-platform
sudo k3s kubectl logs -n octo-dev deploy/octo-orchestrator

Postgres data lives on the boot disk (k3s local-path), so it survives reboots but is destroyed with the VM; the password is held in the release Terraform state. To re-pull a tag manually on the node, sudo octo-pull v0.1.2, which is useful when pods show ImagePullBackOff because the node's registry token (valid ~1h after boot) has expired. To re-bootstrap the VM, sudo rm /opt/octo/.provisioned && sudo reboot. If the release apply cannot reach the cluster, ensure 6443 is open to your IP and refresh the kubeconfig with task deploy:kubeconfig DOMAIN=….

On this page