Octov0.11.7
Deployment

Helm Chart

Install and configure the Octo Helm chart on any Kubernetes cluster.

The octo chart deploys the full platform (the editor UI, the orchestrator, the observability service, Postgres, NATS and Redis) onto any Kubernetes cluster. It is published as an OCI artifact on every release, so installing it needs no checkout:

helm install octo oci://ghcr.io/juancavallotti/charts/octo \--version 0.11.7 \--namespace octo --create-namespace \--set ingress.host=octo.example.com \--set postgres.auth.password=<strong-password>

The defaults point at the public Docker Hub images, and image.tag falls back to the chart's appVersion, so a chart installs the images of its own release rather than drifting onto latest. Both --sets are required: the chart has no default host and no default password, and rendering fails by name without them. If your cluster has no default IngressClass, add --set ingress.className=<yours> as well (see Ingress and TLS).

That command is a demonstration, not an install to keep. --set postgres.auth.password=… puts the credential in Helm's release history for every retained revision. For a deployment you intend to run, follow Install on Kubernetes: the same chart, with every credential in a Secret you own and only its name in the values.

Prefer a cluster-specific profile over the bare defaults. They carry the resource requests, security context and storage class each platform expects. See GKE and EKS.

What the chart deploys

ComponentWorkloadPortNotes
Platform (editor UI)Deployment3000Exposed at ingress.host; proxies the orchestrator and logs APIs through its BFF
OrchestratorDeployment + ServiceAccount + RBAC8090Deploys integrations through the Kubernetes API
Observability serviceDeployment + ServiceAccount + RBAC8091Consumes NATS internal.logs and internal.traces into Postgres; serves the logs, traces, pod stats and retention API. Its Role grants Leases only, for the alerting evaluator
Embeddings serverDeployment8092Vectors for agent-memory search. Off unless embeddings.enabled=true
PostgresStatefulSet + PVC5432Headless Service; credentials in a generated Secret. Skipped when postgres.enabled=false
NATSStatefulSet (no volume)4222 / 8222Core NATS pub-sub; disable with nats.enabled=false
RedisStatefulSet (no volume)6379Shared state between replicas of a component. Replace with a managed one via externalRedis.url; it cannot simply be turned off (see Redis)
Schema applierJob (Helm hook)(none)Runs psql -f schema.sql on every install and upgrade
Retention sweepCronJob(none)Nightly POST /retention/run against the observability service. Disable with retention.enabled=false
Runtime RBACServiceAccount + Role(none)Lease permissions for integration pods (leader election)
Dev-run keySecret(none)HMAC key for dev-run identity and hostnames. Only when orchestrator.devRuns.enabled=true

Five images are not chart workloads. octo-runtime is passed to the orchestrator as RUNTIME_IMAGE, and the orchestrator deploys one pod set per integration from it at runtime. octo-agenticrunner is the alternative for an integration that asked for the agentic runner (see Agentic runner). octo-devruntime and octo-devsidecar are the two containers of a dev run (see Dev runs), created the same way, and octo-statssidecar is injected beside every deployed integration when Pod Stats is on. The schema Job is annotated helm.sh/hook: post-install,post-upgrade with the before-hook-creation delete policy, so it re-runs (idempotently) on every upgrade.

How the chart is put together

octo is an application chart built on octo-common, a library chart holding every reusable template: naming, labels, image references, and the generic Deployment / StatefulSet / Service / RBAC / Secret renderers. The application templates are thin stubs naming a component:

# helm/templates/observability/deployment.yaml (the whole file)
{{- include "octo-common.deployment" (dict "root" $ "component" "observability" ...) }}

So a knob added once applies to every component and every target. Environment variables are the exception, being component-specific, so each component keeps its own _env.tpl. octo-common is vendored at helm/charts/octo-common and packaged inside the octo tarball, so installing octo pulls nothing extra. It is also published standalone at the same version (oci://ghcr.io/juancavallotti/charts/octo-common).

Environment profiles

The chart ships a values file per target. Each is a starting point: you still supply the hostname and a password.

ProfileTargetNotable choices
values-k3d.yamlLocal k3dNodePorts instead of an Ingress, :dev images, no TLS, database pinned to ./data
values-k3d-traefik.yamlLocal k3d + TraefikIngress mode behind the Traefik k3s ships; editor and apps by hostname over HTTPS on port 8443
values-k3d-gateway.yamlLocal k3d + TraefikThe same, through the Gateway API (networking.mode=gateway)
values-kind.yamlLocal kindAs values-k3d.yaml, with kind's port mappings
values-gke-autopilot.yamlGKE AutopilotRequests and limits (Autopilot requires them), restricted security context, generous startup probes
values-gke-standard.yamlGKE StandardRequests without limits so pods can burst, zone spread for the editor
values-eks.yamlAWS EKSalb IngressClass, ACM certificate, gp3 storage

Where to get a profile

The profiles ship inside the published chart, so getting one needs no checkout:

helm pull oci://ghcr.io/juancavallotti/charts/octo --version 0.11.7 --untarls octo/values-*.yaml# octo/values-eks.yaml  octo/values-gke-autopilot.yaml  octo/values-gke-standard.yaml# octo/values-k3d.yaml  octo/values-k3d-gateway.yaml  octo/values-k3d-traefik.yaml  octo/values-kind.yaml

Or download just the one you want:

curl -O https://raw.githubusercontent.com/juancavallotti/octo/main/helm/values-eks.yaml

They are also browsable on GitHub under helm/, where every value carries a comment explaining it. Then install with one:

helm install octo oci://ghcr.io/juancavallotti/charts/octo \--version 0.11.7 -n octo --create-namespace \-f values-eks.yaml --set ingress.host=octo.example.com

Pull the profile at the same version as the chart you install. The profiles pin per-component runAsUser values that match the UIDs that release's images run as.

Values

Defaults live in helm/values.yaml, where every knob is documented inline. The ones that matter most:

Images

KeyDefaultPurpose
image.registrydocker.io/juancavallottiRegistry base the per-component repositories are appended to.
image.tag""Shared tag. Empty falls back to the chart's appVersion, so a chart installs the images of its own release.
image.pullPolicyIfNotPresentApplied to all Octo images.
<component>.image.repositoryocto-<component>-paasPer-component override; <component>.repository still works.
<component>.image.registry(none)Per-component registry, overriding the shared one.
<component>.image.digest(none)Pin by digest. Wins over any tag.

Components are platform, orchestrator, observability, embeddings, schema, runtime, agenticrunner, devsidecar, statssidecar and devruntime. postgres, nats and redis also take image.*, but do not inherit image.registry, which holds the octo-* images only. Point their own image.registry at a mirror for clusters that cannot reach Docker Hub.

The GCP release path pins by digest: Cloud Build renders a values file naming the exact digest of every image it pushed. Render one yourself with task helm:values:images IMAGE_BASE=… TAG=….

Per-workload knobs

Every workload accepts the same set, and defaults: applies them to all of them at once. A per-component value overrides the default; the two are merged, not replaced wholesale.

defaults:
  resources:
    requests: { cpu: 100m, memory: 256Mi }
  securityContext:
    runAsNonRoot: true
    allowPrivilegeEscalation: false
    capabilities: { drop: ["ALL"] }

platform:
  resources:
    requests: { cpu: 250m, memory: 512Mi }   # only this component

Available on defaults and on each component: resources, podSecurityContext, securityContext, imagePullSecrets, nodeSelector, tolerations, affinity, topologySpreadConstraints, priorityClassName, podAnnotations, podLabels, extraEnv, envFrom, extraVolumes, extraVolumeMounts, livenessProbe, startupProbe, revisionHistoryLimit and strategy.

The orchestrator, observability service and runtime additionally take serviceAccount.{create,name,annotations}, which is where GKE Workload Identity and EKS IRSA annotations go.

Private registries

If you mirror the images somewhere that needs credentials, create the pull Secret in the release namespace and name it once:

defaults:
  imagePullSecrets:
    - name: regcred

That reaches all the chart workloads and the integration pods, which the orchestrator creates at runtime from a value passed to it as an environment variable. Without it a mirrored install comes up healthy and every integration you deploy sits in ErrImagePull. Override it for the runtime image alone with runtime.imagePullSecrets; empty there inherits defaults.

Ingress and TLS

ingress.tls.mode is the biggest single difference between targets:

ModeCertificate comes from
cert-managerA ClusterIssuer, issued per host over HTTP-01. What the k3s bootstrap provides.
secretA TLS Secret that already exists: a pre-issued wildcard, or one Terraform placed.
gke-managed-certA Google-managed certificate; the chart creates the ManagedCertificate.
acmAn AWS ACM certificate by ARN, terminated at the ALB.
noneTerminated upstream; no TLS on the Ingress.

Empty auto-selects cert-manager when a clusterIssuer is set, else secret.

Worked examples

Complete values.yaml files, one per mode, each rendering and schema-validating as shown. Install with -f values.yaml.

A ClusterIssuer issues per host over HTTP-01: one certificate for the editor, and one more for each integration you expose.

ingress:
  enabled: true
  className: nginx
  host: octo.example.com
  tls:
    mode: cert-manager
    clusterIssuer: letsencrypt-prod
    secretName: octo-tls

# Per-integration endpoints at {slug}.apps.example.com, each getting its own
# certificate from the same issuer as it is deployed.
orchestrator:
  baseDomain: apps.example.com
  ingressClass: nginx
  clusterIssuer: letsencrypt-prod

postgres:
  auth:
    password: CHANGEME

DNS for every host, including each new subdomain, must already resolve to the ingress address or the HTTP-01 challenge cannot complete. Prefer a wildcard past a handful of hosts.

Where certificates come from

1. Who issues, and who owns the issuer. This chart never creates an Issuer or ClusterIssuer; it names one that must already exist, so clusterIssuer: letsencrypt-prod presumes somebody created it. The k3s bootstrap in this repo does, which is why that name appears in the defaults.

2. How the challenge is solved. For ACME issuers (Let's Encrypt, ZeroSSL) this decides what you can issue:

HTTP-01DNS-01
Proves control byserving a token at http://{host}/.well-known/acme-challenge/…writing a TXT record at _acme-challenge.{host}
Needspublic DNS pointing at the ingress, reachable on port 80API credentials for the DNS zone
Issues wildcardsnoyes
Works on a private clusternoyes
Per new subdomaina fresh challenge, and a waitnothing; the wildcard already covers it

So wildcardTLS.clusterIssuer defaults to a different issuer from ingress.tls.clusterIssuer: *.{baseDomain} needs a DNS solver, while the editor's own host works with HTTP-01.

On a Gateway API cluster with no ingress controller at all, an HTTP-01 issuer has nothing to create its challenge route with unless cert-manager is configured for Gateway API (config.gatewayAPI.enabled=true) so it can use the gatewayHTTPRoute solver. DNS-01 sidesteps that and is the easier answer there.

3. "Cloud certificates" are two unrelated mechanisms. They differ in where the private key lives:

ApproachKey livesSettingcert-manager
Google-managed certificatethe GCE load balancertls.mode: gke-managed-certnot involved
AWS ACMthe ALBtls.mode: acmnot involved
Let's Encrypta Secret in your namespacetls.mode: cert-manageryes
Google CAS / AWS Private CAa Secret in your namespacetls.mode: cert-manager + the annotations belowyes, plus that issuer's own controller

The first two never produce a Kubernetes Secret: the certificate is attached to the load balancer by annotation, which is why those Ingresses carry no tls block and why neither mode has a Gateway API equivalent. The last row is cert-manager issuing from a private CA (a corporate PKI, GoogleCASClusterIssuer, AWSPCAClusterIssuer), which lives outside the cert-manager.io API group, so the reference needs a kind and a group as well as a name, supplied as annotations:

ingress:
  enabled: true
  className: nginx
  host: octo.example.com
  annotations:
    # Out-of-tree issuer. AWS Private CA is:
    #   cert-manager.io/issuer-kind: AWSPCAClusterIssuer
    #   cert-manager.io/issuer-group: awspca.cert-manager.io
    cert-manager.io/issuer-kind: GoogleCASClusterIssuer
    cert-manager.io/issuer-group: cas-issuer.jetstack.io
  tls:
    mode: cert-manager
    clusterIssuer: cas-prod      # the external issuer's name
    secretName: octo-tls

That covers the editor's Ingress, where cert-manager's ingress-shim reads the annotations. The Certificate resources the chart writes itself (the *.{baseDomain} wildcard, and the editor's certificate for a chart-created Gateway) hardcode kind: ClusterIssuer, group: cert-manager.io. For an out-of-tree issuer, set wildcardTLS.clusterIssuer: "" and write the Certificate yourself with the issuerRef you need; the chart then references wildcardTLS.secretName and issues nothing.

ingress.className selects the controller (traefik, nginx, gce, alb); ingress.annotations carries controller-specific behaviour. Per-integration external endpoints are a separate axis: orchestrator.baseDomain plus orchestrator.ingressClass and orchestrator.ingressAnnotations, which the orchestrator stamps onto the Ingresses it generates at runtime.

Both class values default to empty, which omits ingressClassName and lets whichever IngressClass the cluster marks default claim the Ingress. Every environment profile sets it explicitly, and so should you if the cluster has no default, or has one that is not the controller you want:

--set ingress.className=nginx --set orchestrator.ingressClass=nginx

An Ingress that no controller claims is not an error anywhere. It is created, helm install succeeds, and the host never resolves to anything. If the editor is unreachable after an otherwise clean install, check kubectl get ingress for an empty ADDRESS and kubectl get ingressclass for which one, if any, is marked default.

Gateway API

networking.mode selects which Kubernetes API publishes external endpoints, both the editor's and the ones the orchestrator creates per deployed integration. They are served by the same proxy, so it is one switch.

ModeObjectsNeeds
ingressIngress per endpointNothing; built into every cluster. The default.
gatewayHTTPRoute per endpoint, attached to a GatewayThe Gateway API CRDs.

gateway covers kgateway, Envoy Gateway, Istio, Cilium, and the Gateway API modes of ingress-nginx and Traefik.

kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.6.1/standard-install.yamlhelm install octo oci://ghcr.io/juancavallotti/charts/octo \--version 0.11.7 \--namespace octo --create-namespace \--set networking.mode=gateway \--set networking.gateway.name=shared-gw \--set networking.gateway.namespace=ingress \--set ingress.host=octo.example.com

Any Gateway API release from v1.0 onwards works, in the Standard channel; v1.6.1 above is the current one. The orchestrator builds against the gateway.networking.k8s.io/v1 types, GA since v1.0, and uses only Core-support fields (parentRefs, hostnames, rules.matches, rules.backendRefs), so what it sends is accepted unchanged across that range. The Go module sigs.k8s.io/gateway-api is pinned lower than the CRDs you install, which changes nothing: only the serialised shape of an HTTPRoute matters.

Attaching to an existing Gateway is the default. The cluster owner owns the Gateway and octo owns only the routes in its own namespace; the orchestrator's Role grants httproutes and nothing on gateways. The listener owns the ingress class, the certificate arrangement and any controller-specific annotations, so orchestrator.clusterIssuer, orchestrator.ingressClass, orchestrator.ingressAnnotations and wildcardTLS's plumbing have no counterpart under gateway, and the chart stops passing them.

KeyDefaultPurpose
networking.gateway.name""Gateway to attach to. Required unless create is true.
networking.gateway.namespace""Empty is the release namespace. A Gateway elsewhere must allow attachment from here via its listener's allowedRoutes. Cannot be combined with create: a Gateway this chart creates belongs to the release namespace, and the render refuses the pair.
networking.gateway.sectionName""One listener by name. Usually unnecessary, since a route attaches by hostname. Set it where one hostname is served by both an HTTP and an HTTPS listener, or the endpoint is also served in plaintext.
networking.gateway.createfalseHave the chart create the Gateway.
networking.gateway.className""GatewayClass for a created Gateway (kgateway, envoy-gateway, istio, cilium, nginx). Required when create is true.
networking.gateway.annotations{}Extra annotations on a created Gateway.

The Gateway belongs to whoever runs the cluster's ingress; octo adds routes to it and owns nothing else.

networking:
  mode: gateway
  gateway:
    name: shared-gw
    namespace: ingress

ingress:
  enabled: true
  host: octo.example.com
  # className, annotations and tls are Ingress-only and ignored here: the
  # listener on shared-gw decides the controller and holds the certificate.

orchestrator:
  baseDomain: apps.example.com
  # clusterIssuer / ingressClass / ingressAnnotations have no counterpart in
  # this mode; the listener answers all three.

postgres:
  auth:
    password: CHANGEME

shared-gw needs a listener matching each hostname (octo.example.com and *.apps.example.com), and its allowedRoutes.namespaces must permit this release's namespace. The chart cannot set either, so confirm both with the Gateway's owner before installing.

create: true derives one listener per hostname the release already publishes (ingress.host and the *.{baseDomain} wildcard), so a route never has to name one. There is no listener override: anything more elaborate is a Gateway you should own, with create left false.

Certificates. For a Gateway it created, the chart writes the editor's Certificate itself rather than annotating the Gateway for cert-manager's gateway-shim, which needs config.gatewayAPI.enabled=true. On a Gateway you own, its listeners and their certificates are yours, and a Secret in the release namespace could not be referenced from another without a ReferenceGrant. With no ingress controller at all, an HTTP-01 issuer needs cert-manager's gatewayHTTPRoute solver, which is the ClusterIssuer's configuration.

A route that attaches to nothing (a misspelled Gateway, a listener that does not allow your namespace) is not an error anywhere. It is created, helm install succeeds, and the host resolves to nothing. It reports itself in one place:

kubectl -n octo get httproute -o jsonpath='{.items[*].status.parents[*].conditions[*]}'
kubectl -n ingress get gateway shared-gw -o jsonpath='{.status.listeners[*].attachedRoutes}'

Accepted=True on the route and a non-zero attachedRoutes on the listener are what "it worked" looks like.

The orchestrator fails at startup, rather than at the first deploy, when it is configured for Gateway API on a cluster without the CRDs. gke-managed-cert and acm are refused at render time under gateway, having no Gateway API equivalent.

Switching an existing release between modes

The chart's own objects swap over on helm upgrade: the editor's Ingress is replaced by an HTTPRoute, and vice versa. Deployments already running do not. A deployment created under ingress keeps its Ingress after the switch, and undeploying it under gateway removes an HTTPRoute it never had. Redeploy each exposed integration after switching, then sweep up what the old mode left:

# after switching to gateway
kubectl -n octo delete ingress -l app.kubernetes.io/managed-by=orchestrator
# after switching back to ingress
kubectl -n octo delete httproute -l app.kubernetes.io/managed-by=orchestrator

The editor needs a secure context. Parts of it call browser APIs (crypto.randomUUID among them) that exist only over HTTPS or on localhost. Served over plain HTTP on any other hostname the pages still render, but client-side actions throw. tls.mode: none is for endpoints where something upstream terminates TLS, not for serving the editor unencrypted.

Database

The chart runs Postgres itself by default, or points at a managed one.

KeyDefaultPurpose
postgres.enabledtrueDeploy the bundled StatefulSet. false skips it, its Service and its Secret.
postgres.storage.size5GiPVC size.
postgres.storage.storageClassName""Empty uses the cluster default.
postgres.storage.hostPath""Pin the data to a fixed node path. Single-node clusters only; see below.
externalDatabase.host / .port / .database / .user / .sslmode(none)Managed database coordinates, used when postgres.enabled=false.
externalDatabase.existingSecret / .existingSecretPasswordKey(none)Read the password from a Secret you already have.
# Cloud SQL / RDS instead of the bundled StatefulSet
postgres:
  enabled: false
externalDatabase:
  host: octo.abcdef.us-east-1.rds.amazonaws.com
  database: octo
  user: octo
  sslmode: require
  existingSecret: rds-octo
  existingSecretPasswordKey: password

Prefer existingSecret over an inline password, because values persist in Helm release history. Either way the chart passes the password via secretKeyRef and lets Kubernetes expand it into DATABASE_URL, so it is never plaintext in a pod's environment.

Storage on single-node clusters. local-path-style provisioners name each volume's directory after the claim's UID, so recreating the cluster (or the claim) hands Postgres an empty directory while the old one stays on disk. postgres.storage.hostPath pins the data to a fixed path instead, and the chart creates a static PersistentVolume bound to it. Leave it empty on any managed cluster: hostPath is node-local, and GKE Autopilot rejects it. Setting it on a release that already holds data is a data move: copy the directory across with the workload stopped, or the database comes up empty.

Nothing in the chart deletes a database. The StatefulSet sets persistentVolumeClaimRetentionPolicy: Retain, so neither helm uninstall nor scaling to zero removes the claim, and the static volume carries helm.sh/resource-policy: keep.

Secrets and auth

KeyDefaultPurpose
postgres.auth.username / .databaseoctoStored in a chart-managed Secret. The password joins them there only when you supply it inline; with existingSecret it stays in yours.
postgres.auth.password(none)Required. No default: a published chart cannot ship a credential. Rendering fails without it, or without existingSecret below.
postgres.auth.existingSecret / .existingSecretPasswordKey"" / postgres-passwordTake the password from a Secret you already own instead. Preferred for anything long-lived, since every value passed to Helm stays in the release history.
kv.encryptionKey""Base64 32-byte AES-256 key encrypting KV secret namespaces at rest. Empty rejects secret-namespace writes; plain KV still works. See KV and storage.
kv.existingSecret / .existingSecretKey"" / kv-encryption-keyRead that key from a Secret you already own, and the chart creates none. Preferred: this key cannot be rotated to recover from a leak, because a new one makes everything already written to a secret namespace unreadable. Clear kv.encryptionKey when you set it; both together is refused.
auth.oidc.issuer / .clientId(none)Required. IdP details, neither a credential; both are plain env values on the editor. Signing in is the only way into the editor, so rendering fails without them. Any OIDC provider works. Register {auth.url}/api/auth/callback/oidc as its redirect URI.
auth.oidc.clientSecret(none)Required, unless auth.existingSecret below. Lands in the chart's auth Secret.
auth.oidc.providerName / .providerLogo""What the sign-in button calls the provider ("Sign in with …", default OIDC) and the mark beside it (default: the issuer's favicon).
auth.oidc.scopes""Space-separated scopes; empty requests openid profile email.
auth.oidc.endpoints.*""authorization / token / userinfo / jwks overrides, for providers whose discovery document is unusable. Leave empty against a compliant one.
auth.secret""Required, unless auth.existingSecret below. Auth.js session secret (AUTH_SECRET); openssl rand -base64 32 mints one. Rendering fails without it rather than producing a Secret with an empty key, which Auth.js rejects at the first sign-in.
auth.existingSecret""A Secret you already own carrying both credentials, so the chart creates none and requires neither inline value. Clear clientSecret and secret when you set it; both together is refused.
auth.existingSecretClientSecretKey / .existingSecretAuthSecretKeyoidc-client-secret / auth-secretKeys within it. The defaults match the chart's own Secret, so a Secret created with those key names needs neither line.
auth.writeRoles""Roles allowed to mutate. Empty uses every role except platform:monitor.
embeddings.apiKey""Provider key for the embedding server, when one is deployed.
embeddings.existingSecret / .existingSecretKey"" / apiKeyThat key from a Secret you already own, and the chart creates none. Clear apiKey when you set it; both together is refused.
orchestrator.devRuns.hashSecret""Required when dev runs are on. Keys the derivation of every dev run's identity and hostname. See Dev runs for why the chart will not generate one.
orchestrator.devRuns.existingSecret / .existingSecretKey"" / dev-run-hash-secretTake that key from a Secret you own instead. With it set, hashSecret is not needed, and setting both is refused.

Agentic runner

Configures the image the orchestrator deploys for an integration whose deployment asked for runner: agentic: the one carrying a shell, curl, jq, the standalone octo CLI, dolphin and a scratch workspace. Leave these at their defaults unless you are turning the runner off or sizing its pod. The platform agent requires it. See octo-agenticrunner.

KeyDefaultPurpose
agenticrunner.repositoryocto-agenticrunner-paasThe image. Follows the same registry/tag/digest layering as every other component.
agenticrunner.workspaceSize100MiCap on the /workspace emptyDir. A cap, not an allocation: the volume costs nothing until it is written to.
agenticrunner.resources{}Requests and limits for the container, as an ordinary resources block. Passed to the orchestrator as JSON, because integration pods are created by it at deploy time rather than rendered by this chart.

To turn the runner off, clear agenticrunner.repository. The orchestrator then refuses an agentic deploy naming this value, and installing the agent is blocked up front with the reason shown.

Treat the pod as the boundary. A pod holding a shell and a runtime it can point at a definition it just wrote is a general execution environment, so no allow list inside the flow contains it. Everything that pod holds (secrets bound to its environment included) and everything it can reach on the network is reachable by anything it runs. Grant it per deployment on that basis.

Set resources where a pod's appetite matters. No integration pod carries requests or limits by default, and this is the one workload whose purpose is running other programs. GKE Autopilot requires them.

Dev runs

The editor's Run button, executed as a pod the orchestrator creates rather than as a child process of whichever platform replica answered the request. The platform runs several replicas with no session affinity, so a run living in one replica's memory would be invisible to the others. On by default; disabling it makes Run unavailable. See Dev runs for the feature.

KeyDefaultPurpose
orchestrator.devRuns.enabledtrueRenders the two images, the sidecar port, the idle timeout and the key Secret into the orchestrator's environment. Off, the orchestrator reports dev runs unavailable and the editor's Run says so.
orchestrator.devRuns.hashSecret""Required. HMAC key deriving each run's identity and its public hostname from (user, integration).
orchestrator.devRuns.existingSecret / .existingSecretKey"" / dev-run-hash-secretThat key from a Secret you already hold. The requirement is on the pair, not on hashSecret alone.
orchestrator.devRuns.idleTimeout60mHow long a run survives untouched before the orchestrator reaps it. The only bound on how many pods an editing session leaves behind; there is no per-user cap.
orchestrator.devRuns.sidecarPort8099Port the run's sidecar serves its reload/status API on, inside the pod. Never reached from outside it.

The key must never change. A run's public hostname is derived from it, so rotating it re-labels every exposed dev run and a webhook registered against the old hostname silently stops being delivered. The chart requires the value instead of generating one, because a generated key would be regenerated by any render that cannot read the cluster (helm template, --dry-run, a GitOps pipeline). Hold it wherever you hold kv.encryptionKey.

A Secret rather than a values file is the better answer for every credential the chart takes. existingSecret is the same shape everywhere:

postgres:
  auth:
    existingSecret: octo-db-password     # key: postgres-password
kv:
  existingSecret: octo-kv-key            # key: kv-encryption-key
auth:
  existingSecret: octo-auth-creds        # keys: oidc-client-secret, auth-secret
embeddings:
  existingSecret: octo-embeddings-key    # key: apiKey
orchestrator:
  devRuns:
    existingSecret: octo-devrun-key      # key: dev-run-hash-secret

Do not name your Secret after one the chart creates. Those are {release}-postgres, {release}-auth, {release}-kv, {release}-devruns and {release}-embeddings. {release}-postgres collides immediately, since the chart creates it whatever the password's source and Helm refuses to install over a Secret it does not own. The other four collide when you migrate: the upgrade that stops rendering a Secret the previous revision owned deletes the Secret your new revision points at. Neither shows up in helm template.

With references, no key material lands in the release history, and the chart creates no Secret of its own for the auth, KV, dev-run or embeddings credentials. {release}-postgres is still created because it carries the username and database name; with postgres.auth.existingSecret set it holds those two and no password.

--set-file and set_sensitive do not help. They change how a value is supplied and how it is printed, not where it ends up: Helm writes every value it is given into the release Secret in the cluster, and keeps it for every retained revision. A reference is the only form that does not.

Clear the inline value when you switch: setting both is refused, since leaving encryptionKey, hashSecret, apiKey or the two auth values beside an existingSecret would keep the copy the switch was made to avoid. Revisions written before the change still hold what they were given and helm history keeps them, so migrate before anything has been encrypted or deployed.

A dev-run hostname is stable but not private. It is an unguessable hash, not an access control. With orchestrator.clusterIssuer set, each one also gets a per-host certificate and therefore an entry in public CT logs, so prefer wildcardTLS (one *.{baseDomain} certificate) when dev runs are on. Without orchestrator.baseDomain a run still works, reachable only in-cluster.

Data retention

A CronJob that asks the observability service, once a night, to delete stored logs and traces older than the site's retention policy. These values decide whether the job exists and when it runs; how long to keep things is a policy in the database, edited in the admin section. See Data retention.

KeyDefaultPurpose
retention.enabledtrueInstall the CronJob. The policy defaults to keeping everything, so a fresh install sweeps nightly and deletes nothing until somebody sets a window. Turning it off removes the schedule, not the capability: the endpoint stays served, so an on-demand sweep still works.
retention.schedule0 3 * * *When it runs, in the cluster's timezone. The first sweep after a policy is set can be the whole of both tables, and it competes with ingest while it runs.
retention.timeoutSeconds900How long curl waits. It matches the service's own limit on a sweep, so neither side abandons a purge that is making progress.
retention.concurrencyPolicyForbidSkip a run while the previous one is still going, rather than stacking sweeps that would only refuse each other with a 409.
retention.backoffLimit2Retries within one scheduled run. Each batch commits on its own, so a failed sweep is less deleted rather than inconsistent, and tomorrow's run picks up where it stopped.
retention.image.*curlimages/curlThe job is one HTTP POST and the octo images are distroless (no shell, no client). Third-party, so it takes its registry from retention.image.registry alone (empty means Docker Hub) rather than from the shared image.registry, as postgres and nats do.

Scaling

KeyDefaultPurpose
platform.replicas / orchestrator.replicas / observability.replicas1Replica counts. The observability service scales safely: its consumers join NATS queue groups, so replicas compete for messages rather than each storing every one. A retention sweep takes a Postgres advisory lock, so extra replicas cannot sweep concurrently either.
nats.enabledtrueDeploy the NATS broker; when off, live event streams fall back to polling. See Event bus.
redis.enabledtrueDeploy Redis. Unlike NATS this cannot simply be turned off; see Redis.
externalRedis.url(none)A redis:// or rediss:// URL, with no credential in it, for a Redis this chart does not run. Rendered as a literal env value.
externalRedis.existingSecret(none)A Secret you created holding the whole URL, for a server that wants a password. Bound by secretKeyRef, so the credential never enters a workload template. Wins over url.
externalRedis.existingSecretKeyredis-urlThe key within that Secret.
redis.maxmemory256mbThe ceiling the server holds itself to. Keep it below any container memory limit, or the kernel gets there before Redis does and the eviction policy never runs.
redis.maxmemoryPolicyallkeys-lruWhat it does on reaching the ceiling. Evicting the coldest keys costs the folds they held; growing without bound costs the pod.
runtime.servicesModulek8sRuntime-services backend injected into integration pods: k8s (Lease leader election + orchestrator KV), standalone, or api (delegated to a server you implement; set runtime.repository to the octo-api image and inject OCTO_PLATFORM_API_URL too). Cluster deploys want k8s.

Redis

Redis holds state shared between replicas of a component. Today that is one thing: the observability service folds the trace records a streaming block emits (one block.pre-invoke and one block.post-invoke per streamed token) into a single row. Its consumers are a NATS queue group, so one trace's records are spread across replicas.

It is not optional the way NATS is. nats.enabled=false only degrades live event streams to polling, but the aggregator refuses to start without a Redis, so an install with neither the bundled server nor externalRedis.url fails while rendering, with a message naming both values.

Nothing here is persisted: no volume, and the server runs with --save "" and --appendonly no. Losing it costs at most the trace records in the folds that were open. To use a managed Redis instead:

redis:
  enabled: false
externalRedis:
  url: redis://cache.internal:6379

The whole connection is that one string. If it contains a password, put it in a Secret instead. externalRedis.url is rendered into the workload as a literal environment value, readable by anyone who can read workloads:

kubectl create secret generic octo-redis \
  --from-literal=redis-url='rediss://default:pw@cache.example:6380'
redis:
  enabled: false
externalRedis:
  existingSecret: octo-redis   # wins over `url` when both are set
  existingSecretKey: redis-url # the default

The chart will not create that Secret for you. Whether the cluster's Redis is reachable is shown on the Platform services page, alongside Postgres, NATS and the Kubernetes API.

What the chart does not create

Four things a production chart often ships are deliberately absent, each with a workable answer in the meantime:

Not shippedWhyIf you need it
PodDisruptionBudgetA PDB on a single-replica Deployment either permits the eviction it was meant to prevent or blocks node drains outright. It becomes meaningful once you raise replicas, differently per component.Apply your own alongside the release; the components carry standard app.kubernetes.io/component labels to select on.
HorizontalPodAutoscalerThroughput lives in the integration pods the orchestrator creates, not in the chart's own workloads, so an HPA here would scale the wrong thing.Apply your own; platform and observability scale horizontally without coordination, the orchestrator is safe to scale but gains little.
NetworkPolicyCorrect policies depend on what else runs in the namespace and which CNI enforces them, and a policy that is silently unenforced reads as protection that is not there.Write them against the component labels. The orchestrator needs the API server; every service needs Postgres and NATS.
Topology spread / anti-affinity by defaultThe knobs exist (defaults.topologySpreadConstraints, defaults.affinity) and the GKE Standard profile uses them; defaulting them would make single-node clusters unschedulable.Set them per environment; see values-gke-standard.yaml.

Image and chart compatibility

A chart at version X expects the images of release X, which is what it installs by default via the appVersion fallback. This matters because the images run as non-root, and the chart's cloud profiles pin the matching UIDs:

ImageRuns as
octo-platform-paas1000 (node)
octo-orchestrator-paas, octo-observability-paas, octo-embeddings-paas, octo-runtime-paas65532 (distroless nonroot)
octo-schema-paas, postgres70 (postgres)
redis999 (the image's own user)
natsany UID (the profiles pin 1000)

Pinning image.tag to an older release while using a newer chart is untested, and the security contexts in the profiles are the most likely thing to break; relax defaults.securityContext accordingly.

Cluster prerequisites

The chart deploys workloads, not cluster infrastructure. Depending on the values you choose, the cluster must already provide:

  • an ingress controller matching ingress.className; none is installed here
  • cert-manager and a ClusterIssuer, for tls.mode: cert-manager (and for wildcardTLS, which needs a DNS-01 issuer)
  • a StorageClass for the Postgres PVC, unless postgres.storage.hostPath is set or postgres.enabled is false

Upgrade

helm upgrade octo oci://ghcr.io/juancavallotti/charts/octo \--version 0.11.7 --namespace octo --reuse-values

Because image.tag follows appVersion, a chart-version bump moves the images with it. What happens:

  • changed image references rewrite the pod templates, so the platform, orchestrator and observability Deployments roll automatically
  • the schema hook Job re-runs and applies sql/schema.sql; the schema is idempotent (IF NOT EXISTS / ON CONFLICT), so this is safe every time
  • Postgres and its PVC are untouched, so your data survives the upgrade
  • integrations already deployed keep running on the runtime image they were deployed with; redeploy them from the editor to pick up a new octo-runtime

In the GCP reference deployment, Terraform owns the Helm release. Do not run helm upgrade against it by hand; apply the release root (or task deploy TAG=…) instead.

Validating changes

task helm:test lints and renders every profile and pipes the output through kubeconform against the real Kubernetes schemas. CI runs it on every push, so a template change that breaks a target you cannot reach locally still fails on the pull request.

Per-target guides

On this page