Octov0.11.7
Deployment

GKE

Deploy the Octo chart on GKE Autopilot or Standard, with Cloud SQL and Workload Identity.

The chart ships two GKE profiles: values-gke-autopilot.yaml and values-gke-standard.yaml. They describe the same deployment and differ only where Autopilot's managed nodes impose constraints that Standard does not.

Get the profile

Both are shipped inside the published chart, so no checkout is needed:

helm pull oci://ghcr.io/juancavallotti/charts/octo --version 0.11.7 --untarls octo/values-gke-*.yaml# octo/values-gke-autopilot.yaml  octo/values-gke-standard.yaml

Or fetch one directly:

curl -O https://raw.githubusercontent.com/juancavallotti/octo/main/helm/values-gke-autopilot.yaml
curl -O https://raw.githubusercontent.com/juancavallotti/octo/main/helm/values-gke-standard.yaml

Browsable on GitHub: helm/values-gke-autopilot.yaml · helm/values-gke-standard.yaml. Every value is commented with why it is set that way, which is worth reading before you change one.

Install

helm upgrade --install octo oci://ghcr.io/juancavallotti/charts/octo \--version 0.11.7 -n octo --create-namespace \-f values-gke-autopilot.yaml \--set ingress.host=octo.example.com \--set postgres.auth.password=<strong-password>

Pull the profile at the same version as the chart: it pins per-component runAsUser values matching the UIDs that release's images run as.

Both profiles have been installed on live clusters, and both are rendered and schema-checked in CI on every change.

On Autopilot: the wildcard certificate via cert-manager DNS-01, Cloud SQL on a private IP, and an integration deployed from the editor and reached over the wildcard host.

On Standard: the per-host certificate via HTTP-01 and the wildcard via DNS-01, ingress-nginx on the reserved static IP serving the editor over HTTPS, and the bundled Postgres StatefulSet on a premium-rwo volume with the schema Job completing against it.

Between them that covers both database paths. The two Terraform roots are thin wrappers over one shared module, and nothing in the Cloud SQL path branches on Autopilot versus Standard. Neither values profile carries an externalDatabase block at all, since that wiring comes from the module.

Provision a cluster with Terraform

If you do not already have a cluster, the repo has a Terraform root per flavour that creates one, installs every prerequisite below, wires DNS, and deploys the chart, in a single apply.

Each root reads its own terraform.tfvars, with nothing shared between them, so copy the example for the flavour you are applying and fill in project_id, acme_email, dns_managed_zone, domain and apps_domain:

cd deploy/terraform

# GKE Standard
cp gke-standard/terraform.tfvars.example gke-standard/terraform.tfvars
$EDITOR gke-standard/terraform.tfvars
task gke:standard:apply

# ...or GKE Autopilot
cp gke-autopilot/terraform.tfvars.example gke-autopilot/terraform.tfvars
$EDITOR gke-autopilot/terraform.tfvars
task gke:autopilot:apply

Both roots default to the same hostnames, so only one can hold them at a time: destroy one before applying the other, or the second apply fails on an existing record set.

Both roots are thin wrappers over one shared composition (deploy/terraform/modules/octo-gke), so the two flavours cannot drift. The point of having both is that the same deployment is exercised on each. What they create:

ClusterA dedicated VPC (not default), VPC-native, Autopilot or Standard
Ingressingress-nginx on a reserved static IP, so DNS resolves before the controller is even up
TLScert-manager, plus both ClusterIssuers: HTTP-01 for the editor host, DNS-01 (via Workload Identity) for the wildcard
DNSA records for the editor, the apps domain and *.{apps_domain}
DatabaseThe chart's bundled Postgres, or a Cloud SQL instance on a private IP, selected by one variable

These are test environments and their defaults say so: spot nodes, a zonal control plane, no Cloud NAT, no deletion protection, no database backups. Every one has a variable that turns it into the production choice.

Treat these roots as a worked example, not as your production infrastructure. One root owns the cluster and the application release, so every terraform apply is also a redeploy; there is a single environment with no staging/production split; the VPC, IAM and state layout are sized for something you will destroy this afternoon.

The chart is the stable interface here, not the Terraform: none of what these roots do to it (hostnames, TLS mode, external database, secrets) requires their Terraform.

The DNS zone is not managed by Terraform, only the records in it. Cloud DNS assigns fresh nameservers on every zone creation, so a zone owned by a root you destroy and recreate would invalidate its delegation every cycle. Create it once by hand and delegate it. See deploy/terraform/README.md.

Tear down with task gke:standard:destroy. It is phased on purpose: the ingress controller owns the load balancer, and destroying it alongside the release strands a forwarding rule that then blocks the VPC delete.

The same target also sweeps orphaned k8s-* firewall rules off the VPC before deleting it. A LoadBalancer Service makes the GKE service controller create five objects Terraform never sees (a forwarding rule, a target pool, a health check and two firewall rules); deleting the Service reclaims them, but that cleanup races the cluster delete, and the health-check rule loses often enough to matter. A network cannot be deleted while a firewall rule references it, so without the sweep the destroy fails on the VPC with an error naming the rule rather than the cause:

Error waiting for Deleting Network: The network resource '…/networks/octo-gke-std'
is already being used by '…/firewalls/k8s-8d299831b5154b81-node-http-hc'

Should you ever hit that, from a destroy interrupted before the sweep or a cluster torn down by hand, it clears with:

gcloud compute firewall-rules list --project <project> --filter='name~^k8s-'
gcloud compute firewall-rules delete <name> --project <project> --quiet

Then re-run the destroy; it is safe to repeat.

Testing a chart you have edited

The published images track releases, so they will not match a chart changed since the last one. Build the working tree instead and push it to ttl.sh:

task images:ttl TTL=12h
task gke:standard:apply IMAGE_VALUES=dist/values.ttl.yaml

Prerequisites

None of these are installed by the chart:

RequirementWhy
An ingress controller (ingress-nginx is what the profiles assume)The chart creates an Ingress; nothing serves it otherwise
cert-manager plus the ClusterIssuer named in ingress.tls.clusterIssuerTLS for the editor host
A DNS A record for ingress.host pointing at the controller's address(none)

The Terraform roots above install all three. This section is for a cluster you already own.

Add a wildcard record for *.{orchestrator.baseDomain} if you want per-integration external endpoints.

Why ingress-nginx rather than GCE ingress

The orchestrator creates an Ingress per externally-exposed integration, at runtime, in Go. GCE ingress requires a cloud.google.com/neg annotation on each backend Service, which the orchestrator does not set, and a Google-managed certificate is a separate CRD it cannot emit. ingress-nginx works with the ClusterIP Services it does create, unchanged.

You can still put the editor behind GCE ingress while integrations use nginx:

ingress:
  className: gce
  tls:
    mode: gke-managed-cert     # the chart creates the ManagedCertificate
orchestrator:
  ingressClass: nginx          # leave per-integration endpoints on nginx

In that configuration you will likely also want gke.backendConfig.enabled (the GCE load balancer health-checks / by default and marks a backend unhealthy if the app answers otherwise) and gke.frontendConfig.enabled for an HTTP→HTTPS redirect. Both need GKE-only CRDs, so both stay off unless you enable them.

Autopilot's constraints

Autopilot manages nodes for you, and three consequences shape the profile.

Every pod is billed on its resource requests, and Autopilot injects defaults when they are absent, so the profile sets them explicitly with limits equal to requests: the Guaranteed QoS class, and a predictable bill.

hostPath and hostPort are rejected outright. Nothing in the chart uses them by default; leave postgres.storage.hostPath empty, as it is a single-node-cluster knob.

Pods can take noticeably longer to start, because Autopilot may provision a node first. The profile allows a generous startup probe (30 × 10s) so a slow first boot is not killed before it is ready.

values-gke-standard.yaml relaxes all three: requests without limits so pods can burst into spare node capacity, no startup-probe allowance, and nodeSelector / tolerations available if you run a dedicated pool. It also spreads the editor across zones so a zonal outage does not take it down.

Security context

Every container runs unprivileged: runAsNonRoot, allowPrivilegeEscalation: false, all capabilities dropped, and seccompProfile: RuntimeDefault. Because each image runs as a different UID, runAsUser is set per component. See image and chart compatibility for the table.

The Go services (orchestrator, observability) additionally get readOnlyRootFilesystem: true, since they write nothing to disk. The editor does not, because Next.js needs writable scratch space.

Storage

AutopilotStandard
storageClassNamestandard-rwo (pd-balanced)premium-rwo (pd-ssd)
Size20Gi50Gi

Persistent disks are zonal, so the Postgres pod is pinned to its volume's zone. That is one of the reasons to prefer Cloud SQL for anything you care about.

Cloud SQL

postgres:
  enabled: false
externalDatabase:
  host: 10.20.30.40          # private IP, or a Cloud SQL Auth proxy
  database: octo
  user: octo
  sslmode: require
  existingSecret: cloudsql-octo
  existingSecretPasswordKey: password

postgres.enabled: false skips the StatefulSet, its Service and its Secret; the schema Job points at the external host instead. Keep the password in a Secret owned by Terraform or external-secrets rather than in a values file, because values persist in Helm release history.

Workload Identity

Bind the chart's ServiceAccounts to Google service accounts through annotations:

orchestrator:
  serviceAccount:
    annotations:
      iam.gke.io/gcp-service-account: octo-orchestrator@PROJECT.iam.gserviceaccount.com

The orchestrator and runtime ServiceAccounts are the two the chart creates and the two that can need cloud credentials.

A note on the GCP reference deployment

This page is about installing the chart on a GKE cluster you own. It is not the same thing as the GCP reference deployment, which provisions a single-VM k3s cluster with Terraform and drives releases through Cloud Build. That path is more opinionated and more automated; this one assumes you already have a cluster and want the platform on it.

On this page