Helm Chart
Install and configure the Octo Helm chart on any Kubernetes cluster.
The octo chart deploys the full platform (the editor UI, the orchestrator,
the observability service, Postgres, NATS and Redis) onto any Kubernetes
cluster. It is published as an OCI artifact on every release, so installing it
needs no checkout:
helm install octo oci://ghcr.io/juancavallotti/charts/octo \--version 0.11.7 \--namespace octo --create-namespace \--set ingress.host=octo.example.com \--set postgres.auth.password=<strong-password>The defaults point at the public Docker Hub images, and image.tag falls back
to the chart's appVersion, so a chart installs the images of its own release
rather than drifting onto latest. Both --sets are required: the chart has no
default host and no default password, and rendering fails by name without them.
If your cluster has no default IngressClass, add
--set ingress.className=<yours> as well (see Ingress and TLS).
That command is a demonstration, not an install to keep. --set postgres.auth.password=… puts the credential in Helm's release history for
every retained revision. For a deployment you intend to run, follow
Install on Kubernetes: the same chart, with every
credential in a Secret you own and only its name in the values.
What the chart deploys
| Component | Workload | Port | Notes |
|---|---|---|---|
| Platform (editor UI) | Deployment | 3000 | Exposed at ingress.host; proxies the orchestrator and logs APIs through its BFF |
| Orchestrator | Deployment + ServiceAccount + RBAC | 8090 | Deploys integrations through the Kubernetes API |
| Observability service | Deployment + ServiceAccount + RBAC | 8091 | Consumes NATS internal.logs and internal.traces into Postgres; serves the logs, traces, pod stats and retention API. Its Role grants Leases only, for the alerting evaluator |
| Embeddings server | Deployment | 8092 | Vectors for agent-memory search. Off unless embeddings.enabled=true |
| Postgres | StatefulSet + PVC | 5432 | Headless Service; credentials in a generated Secret. Skipped when postgres.enabled=false |
| NATS | StatefulSet (no volume) | 4222 / 8222 | Core NATS pub-sub; disable with nats.enabled=false |
| Redis | StatefulSet (no volume) | 6379 | Shared state between replicas of a component. Replace with a managed one via externalRedis.url; it cannot simply be turned off (see Redis) |
| Schema applier | Job (Helm hook) | (none) | Runs psql -f schema.sql on every install and upgrade |
| Retention sweep | CronJob | (none) | Nightly POST /retention/run against the observability service. Disable with retention.enabled=false |
| Runtime RBAC | ServiceAccount + Role | (none) | Lease permissions for integration pods (leader election) |
| Dev-run key | Secret | (none) | HMAC key for dev-run identity and hostnames. Only when orchestrator.devRuns.enabled=true |
Five images are not chart workloads. octo-runtime is passed to the
orchestrator as RUNTIME_IMAGE, and the orchestrator deploys one pod set per
integration from it at runtime. octo-agenticrunner is the alternative for an
integration that asked for the agentic runner (see
Agentic runner). octo-devruntime and octo-devsidecar are
the two containers of a dev run (see Dev runs), created the same
way, and octo-statssidecar is injected beside every deployed integration when
Pod Stats is on. The schema Job is annotated
helm.sh/hook: post-install,post-upgrade with the before-hook-creation delete
policy, so it re-runs (idempotently) on every upgrade.
How the chart is put together
octo is an application chart built on octo-common, a library chart
holding every reusable template: naming, labels, image references, and the
generic Deployment / StatefulSet / Service / RBAC / Secret renderers. The
application templates are thin stubs naming a component:
# helm/templates/observability/deployment.yaml (the whole file)
{{- include "octo-common.deployment" (dict "root" $ "component" "observability" ...) }}So a knob added once applies to every component and every target. Environment
variables are the exception, being component-specific, so each component keeps
its own _env.tpl. octo-common is vendored at helm/charts/octo-common and
packaged inside the octo tarball, so installing octo pulls nothing extra. It
is also published standalone at the same version
(oci://ghcr.io/juancavallotti/charts/octo-common).
Environment profiles
The chart ships a values file per target. Each is a starting point: you still supply the hostname and a password.
| Profile | Target | Notable choices |
|---|---|---|
values-k3d.yaml | Local k3d | NodePorts instead of an Ingress, :dev images, no TLS, database pinned to ./data |
values-k3d-traefik.yaml | Local k3d + Traefik | Ingress mode behind the Traefik k3s ships; editor and apps by hostname over HTTPS on port 8443 |
values-k3d-gateway.yaml | Local k3d + Traefik | The same, through the Gateway API (networking.mode=gateway) |
values-kind.yaml | Local kind | As values-k3d.yaml, with kind's port mappings |
values-gke-autopilot.yaml | GKE Autopilot | Requests and limits (Autopilot requires them), restricted security context, generous startup probes |
values-gke-standard.yaml | GKE Standard | Requests without limits so pods can burst, zone spread for the editor |
values-eks.yaml | AWS EKS | alb IngressClass, ACM certificate, gp3 storage |
Where to get a profile
The profiles ship inside the published chart, so getting one needs no checkout:
helm pull oci://ghcr.io/juancavallotti/charts/octo --version 0.11.7 --untarls octo/values-*.yaml# octo/values-eks.yaml octo/values-gke-autopilot.yaml octo/values-gke-standard.yaml# octo/values-k3d.yaml octo/values-k3d-gateway.yaml octo/values-k3d-traefik.yaml octo/values-kind.yamlOr download just the one you want:
curl -O https://raw.githubusercontent.com/juancavallotti/octo/main/helm/values-eks.yamlThey are also browsable on GitHub under
helm/, where every
value carries a comment explaining it. Then install with one:
helm install octo oci://ghcr.io/juancavallotti/charts/octo \--version 0.11.7 -n octo --create-namespace \-f values-eks.yaml --set ingress.host=octo.example.comPull the profile at the same version as the chart you install. The
profiles pin per-component runAsUser values that match the UIDs that
release's images run as.
Values
Defaults live in helm/values.yaml, where every knob is documented inline. The
ones that matter most:
Images
| Key | Default | Purpose |
|---|---|---|
image.registry | docker.io/juancavallotti | Registry base the per-component repositories are appended to. |
image.tag | "" | Shared tag. Empty falls back to the chart's appVersion, so a chart installs the images of its own release. |
image.pullPolicy | IfNotPresent | Applied to all Octo images. |
<component>.image.repository | octo-<component>-paas | Per-component override; <component>.repository still works. |
<component>.image.registry | (none) | Per-component registry, overriding the shared one. |
<component>.image.digest | (none) | Pin by digest. Wins over any tag. |
Components are platform, orchestrator, observability, embeddings,
schema, runtime, agenticrunner, devsidecar, statssidecar and
devruntime. postgres, nats and redis also take image.*, but do not
inherit image.registry, which holds the octo-* images only. Point their own
image.registry at a mirror for clusters that cannot reach Docker Hub.
The GCP release path pins by digest: Cloud Build renders a values file naming
the exact digest of every image it pushed. Render one yourself with
task helm:values:images IMAGE_BASE=… TAG=….
Per-workload knobs
Every workload accepts the same set, and defaults: applies them to all of
them at once. A per-component value overrides the default; the two are merged,
not replaced wholesale.
defaults:
resources:
requests: { cpu: 100m, memory: 256Mi }
securityContext:
runAsNonRoot: true
allowPrivilegeEscalation: false
capabilities: { drop: ["ALL"] }
platform:
resources:
requests: { cpu: 250m, memory: 512Mi } # only this componentAvailable on defaults and on each component: resources,
podSecurityContext, securityContext, imagePullSecrets, nodeSelector,
tolerations, affinity, topologySpreadConstraints, priorityClassName,
podAnnotations, podLabels, extraEnv, envFrom, extraVolumes,
extraVolumeMounts, livenessProbe, startupProbe, revisionHistoryLimit
and strategy.
The orchestrator, observability service and runtime additionally take
serviceAccount.{create,name,annotations}, which is where GKE Workload Identity
and EKS IRSA annotations go.
Private registries
If you mirror the images somewhere that needs credentials, create the pull Secret in the release namespace and name it once:
defaults:
imagePullSecrets:
- name: regcredThat reaches all the chart workloads and the integration pods, which the
orchestrator creates at runtime from a value passed to it as an environment
variable. Without it a mirrored install comes up healthy and every integration
you deploy sits in ErrImagePull. Override it for the runtime image alone with
runtime.imagePullSecrets; empty there inherits defaults.
Ingress and TLS
ingress.tls.mode is the biggest single difference between targets:
| Mode | Certificate comes from |
|---|---|
cert-manager | A ClusterIssuer, issued per host over HTTP-01. What the k3s bootstrap provides. |
secret | A TLS Secret that already exists: a pre-issued wildcard, or one Terraform placed. |
gke-managed-cert | A Google-managed certificate; the chart creates the ManagedCertificate. |
acm | An AWS ACM certificate by ARN, terminated at the ALB. |
none | Terminated upstream; no TLS on the Ingress. |
Empty auto-selects cert-manager when a clusterIssuer is set, else secret.
Worked examples
Complete values.yaml files, one per mode, each rendering and schema-validating
as shown. Install with -f values.yaml.
A ClusterIssuer issues per host over HTTP-01: one certificate for the editor, and one more for each integration you expose.
ingress:
enabled: true
className: nginx
host: octo.example.com
tls:
mode: cert-manager
clusterIssuer: letsencrypt-prod
secretName: octo-tls
# Per-integration endpoints at {slug}.apps.example.com, each getting its own
# certificate from the same issuer as it is deployed.
orchestrator:
baseDomain: apps.example.com
ingressClass: nginx
clusterIssuer: letsencrypt-prod
postgres:
auth:
password: CHANGEMEDNS for every host, including each new subdomain, must already resolve to the ingress address or the HTTP-01 challenge cannot complete. Prefer a wildcard past a handful of hosts.
Where certificates come from
1. Who issues, and who owns the issuer. This chart never creates an Issuer
or ClusterIssuer; it names one that must already exist, so
clusterIssuer: letsencrypt-prod presumes somebody created it. The k3s bootstrap
in this repo does, which is why that name appears in the defaults.
2. How the challenge is solved. For ACME issuers (Let's Encrypt, ZeroSSL) this decides what you can issue:
| HTTP-01 | DNS-01 | |
|---|---|---|
| Proves control by | serving a token at http://{host}/.well-known/acme-challenge/… | writing a TXT record at _acme-challenge.{host} |
| Needs | public DNS pointing at the ingress, reachable on port 80 | API credentials for the DNS zone |
| Issues wildcards | no | yes |
| Works on a private cluster | no | yes |
| Per new subdomain | a fresh challenge, and a wait | nothing; the wildcard already covers it |
So wildcardTLS.clusterIssuer defaults to a different issuer from
ingress.tls.clusterIssuer: *.{baseDomain} needs a DNS solver, while the
editor's own host works with HTTP-01.
On a Gateway API cluster with no ingress controller at all, an HTTP-01 issuer
has nothing to create its challenge route with unless cert-manager is
configured for Gateway API (config.gatewayAPI.enabled=true) so it can use the
gatewayHTTPRoute solver. DNS-01 sidesteps that and is the easier answer
there.
3. "Cloud certificates" are two unrelated mechanisms. They differ in where the private key lives:
| Approach | Key lives | Setting | cert-manager |
|---|---|---|---|
| Google-managed certificate | the GCE load balancer | tls.mode: gke-managed-cert | not involved |
| AWS ACM | the ALB | tls.mode: acm | not involved |
| Let's Encrypt | a Secret in your namespace | tls.mode: cert-manager | yes |
| Google CAS / AWS Private CA | a Secret in your namespace | tls.mode: cert-manager + the annotations below | yes, plus that issuer's own controller |
The first two never produce a Kubernetes Secret: the certificate is attached to
the load balancer by annotation, which is why those Ingresses carry no tls
block and why neither mode has a Gateway API equivalent. The last row is
cert-manager issuing from a private CA (a corporate PKI,
GoogleCASClusterIssuer, AWSPCAClusterIssuer), which lives outside the
cert-manager.io API group, so the reference needs a kind and a group as well as
a name, supplied as annotations:
ingress:
enabled: true
className: nginx
host: octo.example.com
annotations:
# Out-of-tree issuer. AWS Private CA is:
# cert-manager.io/issuer-kind: AWSPCAClusterIssuer
# cert-manager.io/issuer-group: awspca.cert-manager.io
cert-manager.io/issuer-kind: GoogleCASClusterIssuer
cert-manager.io/issuer-group: cas-issuer.jetstack.io
tls:
mode: cert-manager
clusterIssuer: cas-prod # the external issuer's name
secretName: octo-tlsThat covers the editor's Ingress, where cert-manager's ingress-shim reads the
annotations. The Certificate resources the chart writes itself (the
*.{baseDomain} wildcard, and the editor's certificate for a chart-created
Gateway) hardcode kind: ClusterIssuer, group: cert-manager.io. For an
out-of-tree issuer, set wildcardTLS.clusterIssuer: "" and write the
Certificate yourself with the issuerRef you need; the chart then references
wildcardTLS.secretName and issues nothing.
ingress.className selects the controller (traefik, nginx, gce, alb);
ingress.annotations carries controller-specific behaviour. Per-integration
external endpoints are a separate axis: orchestrator.baseDomain plus
orchestrator.ingressClass and orchestrator.ingressAnnotations, which the
orchestrator stamps onto the Ingresses it generates at runtime.
Both class values default to empty, which omits ingressClassName and lets
whichever IngressClass the cluster marks default claim the Ingress. Every
environment profile sets it explicitly, and so should you if the cluster has no
default, or has one that is not the controller you want:
--set ingress.className=nginx --set orchestrator.ingressClass=nginxAn Ingress that no controller claims is not an error anywhere. It is created,
helm install succeeds, and the host never resolves to anything. If the editor
is unreachable after an otherwise clean install, check kubectl get ingress
for an empty ADDRESS and kubectl get ingressclass for which one, if any, is
marked default.
Gateway API
networking.mode selects which Kubernetes API publishes external endpoints,
both the editor's and the ones the orchestrator creates per deployed integration.
They are served by the same proxy, so it is one switch.
| Mode | Objects | Needs |
|---|---|---|
ingress | Ingress per endpoint | Nothing; built into every cluster. The default. |
gateway | HTTPRoute per endpoint, attached to a Gateway | The Gateway API CRDs. |
gateway covers kgateway, Envoy Gateway, Istio, Cilium, and the Gateway API
modes of ingress-nginx and Traefik.
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.6.1/standard-install.yamlhelm install octo oci://ghcr.io/juancavallotti/charts/octo \--version 0.11.7 \--namespace octo --create-namespace \--set networking.mode=gateway \--set networking.gateway.name=shared-gw \--set networking.gateway.namespace=ingress \--set ingress.host=octo.example.comAny Gateway API release from v1.0 onwards works, in the Standard channel;
v1.6.1 above is the current one. The orchestrator builds against the
gateway.networking.k8s.io/v1 types, GA since v1.0, and uses only Core-support
fields (parentRefs, hostnames, rules.matches, rules.backendRefs), so what
it sends is accepted unchanged across that range. The Go module
sigs.k8s.io/gateway-api is pinned lower than the CRDs you install, which
changes nothing: only the serialised shape of an HTTPRoute matters.
Attaching to an existing Gateway is the default. The cluster owner owns the
Gateway and octo owns only the routes in its own namespace; the orchestrator's
Role grants httproutes and nothing on gateways. The listener owns the ingress
class, the certificate arrangement and any controller-specific annotations, so
orchestrator.clusterIssuer, orchestrator.ingressClass,
orchestrator.ingressAnnotations and wildcardTLS's plumbing have no
counterpart under gateway, and the chart stops passing them.
| Key | Default | Purpose |
|---|---|---|
networking.gateway.name | "" | Gateway to attach to. Required unless create is true. |
networking.gateway.namespace | "" | Empty is the release namespace. A Gateway elsewhere must allow attachment from here via its listener's allowedRoutes. Cannot be combined with create: a Gateway this chart creates belongs to the release namespace, and the render refuses the pair. |
networking.gateway.sectionName | "" | One listener by name. Usually unnecessary, since a route attaches by hostname. Set it where one hostname is served by both an HTTP and an HTTPS listener, or the endpoint is also served in plaintext. |
networking.gateway.create | false | Have the chart create the Gateway. |
networking.gateway.className | "" | GatewayClass for a created Gateway (kgateway, envoy-gateway, istio, cilium, nginx). Required when create is true. |
networking.gateway.annotations | {} | Extra annotations on a created Gateway. |
The Gateway belongs to whoever runs the cluster's ingress; octo adds routes to it and owns nothing else.
networking:
mode: gateway
gateway:
name: shared-gw
namespace: ingress
ingress:
enabled: true
host: octo.example.com
# className, annotations and tls are Ingress-only and ignored here: the
# listener on shared-gw decides the controller and holds the certificate.
orchestrator:
baseDomain: apps.example.com
# clusterIssuer / ingressClass / ingressAnnotations have no counterpart in
# this mode; the listener answers all three.
postgres:
auth:
password: CHANGEMEshared-gw needs a listener matching each hostname (octo.example.com and
*.apps.example.com), and its allowedRoutes.namespaces must permit this
release's namespace. The chart cannot set either, so confirm both with the
Gateway's owner before installing.
create: true derives one listener per hostname the release already publishes
(ingress.host and the *.{baseDomain} wildcard), so a route never has to name
one. There is no listener override: anything more elaborate is a Gateway you
should own, with create left false.
Certificates. For a Gateway it created, the chart writes the editor's
Certificate itself rather than annotating the Gateway for cert-manager's
gateway-shim, which needs config.gatewayAPI.enabled=true. On a Gateway you own,
its listeners and their certificates are yours, and a Secret in the release
namespace could not be referenced from another without a ReferenceGrant. With
no ingress controller at all, an HTTP-01 issuer needs cert-manager's
gatewayHTTPRoute solver, which is the ClusterIssuer's configuration.
A route that attaches to nothing (a misspelled Gateway, a listener that does
not allow your namespace) is not an error anywhere. It is created,
helm install succeeds, and the host resolves to nothing. It reports itself
in one place:
kubectl -n octo get httproute -o jsonpath='{.items[*].status.parents[*].conditions[*]}'
kubectl -n ingress get gateway shared-gw -o jsonpath='{.status.listeners[*].attachedRoutes}'Accepted=True on the route and a non-zero attachedRoutes on the listener
are what "it worked" looks like.
The orchestrator fails at startup, rather than at the first deploy, when it is
configured for Gateway API on a cluster without the CRDs. gke-managed-cert and
acm are refused at render time under gateway, having no Gateway API
equivalent.
Switching an existing release between modes
The chart's own objects swap over on helm upgrade: the editor's Ingress is
replaced by an HTTPRoute, and vice versa. Deployments already running do not.
A deployment created under ingress keeps its Ingress after the switch, and
undeploying it under gateway removes an HTTPRoute it never had. Redeploy each
exposed integration after switching, then sweep up what the old mode left:
# after switching to gateway
kubectl -n octo delete ingress -l app.kubernetes.io/managed-by=orchestrator
# after switching back to ingress
kubectl -n octo delete httproute -l app.kubernetes.io/managed-by=orchestratorThe editor needs a secure context. Parts of it call browser APIs
(crypto.randomUUID among them) that exist only over HTTPS or on localhost.
Served over plain HTTP on any other hostname the pages still render, but
client-side actions throw. tls.mode: none is for endpoints where something
upstream terminates TLS, not for serving the editor unencrypted.
Database
The chart runs Postgres itself by default, or points at a managed one.
| Key | Default | Purpose |
|---|---|---|
postgres.enabled | true | Deploy the bundled StatefulSet. false skips it, its Service and its Secret. |
postgres.storage.size | 5Gi | PVC size. |
postgres.storage.storageClassName | "" | Empty uses the cluster default. |
postgres.storage.hostPath | "" | Pin the data to a fixed node path. Single-node clusters only; see below. |
externalDatabase.host / .port / .database / .user / .sslmode | (none) | Managed database coordinates, used when postgres.enabled=false. |
externalDatabase.existingSecret / .existingSecretPasswordKey | (none) | Read the password from a Secret you already have. |
# Cloud SQL / RDS instead of the bundled StatefulSet
postgres:
enabled: false
externalDatabase:
host: octo.abcdef.us-east-1.rds.amazonaws.com
database: octo
user: octo
sslmode: require
existingSecret: rds-octo
existingSecretPasswordKey: passwordPrefer existingSecret over an inline password, because values persist in Helm
release history. Either way the chart passes the password via secretKeyRef and
lets Kubernetes expand it into DATABASE_URL, so it is never plaintext in a
pod's environment.
Storage on single-node clusters. local-path-style provisioners name each
volume's directory after the claim's UID, so recreating the cluster (or the
claim) hands Postgres an empty directory while the old one stays on disk.
postgres.storage.hostPath pins the data to a fixed path instead, and the
chart creates a static PersistentVolume bound to it. Leave it empty on any
managed cluster: hostPath is node-local, and GKE Autopilot rejects it. Setting
it on a release that already holds data is a data move: copy the directory
across with the workload stopped, or the database comes up empty.
Nothing in the chart deletes a database. The StatefulSet sets
persistentVolumeClaimRetentionPolicy: Retain, so neither helm uninstall nor
scaling to zero removes the claim, and the static volume carries
helm.sh/resource-policy: keep.
Secrets and auth
| Key | Default | Purpose |
|---|---|---|
postgres.auth.username / .database | octo | Stored in a chart-managed Secret. The password joins them there only when you supply it inline; with existingSecret it stays in yours. |
postgres.auth.password | (none) | Required. No default: a published chart cannot ship a credential. Rendering fails without it, or without existingSecret below. |
postgres.auth.existingSecret / .existingSecretPasswordKey | "" / postgres-password | Take the password from a Secret you already own instead. Preferred for anything long-lived, since every value passed to Helm stays in the release history. |
kv.encryptionKey | "" | Base64 32-byte AES-256 key encrypting KV secret namespaces at rest. Empty rejects secret-namespace writes; plain KV still works. See KV and storage. |
kv.existingSecret / .existingSecretKey | "" / kv-encryption-key | Read that key from a Secret you already own, and the chart creates none. Preferred: this key cannot be rotated to recover from a leak, because a new one makes everything already written to a secret namespace unreadable. Clear kv.encryptionKey when you set it; both together is refused. |
auth.oidc.issuer / .clientId | (none) | Required. IdP details, neither a credential; both are plain env values on the editor. Signing in is the only way into the editor, so rendering fails without them. Any OIDC provider works. Register {auth.url}/api/auth/callback/oidc as its redirect URI. |
auth.oidc.clientSecret | (none) | Required, unless auth.existingSecret below. Lands in the chart's auth Secret. |
auth.oidc.providerName / .providerLogo | "" | What the sign-in button calls the provider ("Sign in with …", default OIDC) and the mark beside it (default: the issuer's favicon). |
auth.oidc.scopes | "" | Space-separated scopes; empty requests openid profile email. |
auth.oidc.endpoints.* | "" | authorization / token / userinfo / jwks overrides, for providers whose discovery document is unusable. Leave empty against a compliant one. |
auth.secret | "" | Required, unless auth.existingSecret below. Auth.js session secret (AUTH_SECRET); openssl rand -base64 32 mints one. Rendering fails without it rather than producing a Secret with an empty key, which Auth.js rejects at the first sign-in. |
auth.existingSecret | "" | A Secret you already own carrying both credentials, so the chart creates none and requires neither inline value. Clear clientSecret and secret when you set it; both together is refused. |
auth.existingSecretClientSecretKey / .existingSecretAuthSecretKey | oidc-client-secret / auth-secret | Keys within it. The defaults match the chart's own Secret, so a Secret created with those key names needs neither line. |
auth.writeRoles | "" | Roles allowed to mutate. Empty uses every role except platform:monitor. |
embeddings.apiKey | "" | Provider key for the embedding server, when one is deployed. |
embeddings.existingSecret / .existingSecretKey | "" / apiKey | That key from a Secret you already own, and the chart creates none. Clear apiKey when you set it; both together is refused. |
orchestrator.devRuns.hashSecret | "" | Required when dev runs are on. Keys the derivation of every dev run's identity and hostname. See Dev runs for why the chart will not generate one. |
orchestrator.devRuns.existingSecret / .existingSecretKey | "" / dev-run-hash-secret | Take that key from a Secret you own instead. With it set, hashSecret is not needed, and setting both is refused. |
Agentic runner
Configures the image the orchestrator deploys for an integration whose
deployment asked for runner: agentic: the one carrying a shell, curl, jq,
the standalone octo CLI, dolphin and a scratch workspace. Leave these at
their defaults unless you are turning the runner off or sizing its pod. The
platform agent requires it. See
octo-agenticrunner.
| Key | Default | Purpose |
|---|---|---|
agenticrunner.repository | octo-agenticrunner-paas | The image. Follows the same registry/tag/digest layering as every other component. |
agenticrunner.workspaceSize | 100Mi | Cap on the /workspace emptyDir. A cap, not an allocation: the volume costs nothing until it is written to. |
agenticrunner.resources | {} | Requests and limits for the container, as an ordinary resources block. Passed to the orchestrator as JSON, because integration pods are created by it at deploy time rather than rendered by this chart. |
To turn the runner off, clear agenticrunner.repository. The orchestrator
then refuses an agentic deploy naming this value, and installing the agent is
blocked up front with the reason shown.
Treat the pod as the boundary. A pod holding a shell and a runtime it can point at a definition it just wrote is a general execution environment, so no allow list inside the flow contains it. Everything that pod holds (secrets bound to its environment included) and everything it can reach on the network is reachable by anything it runs. Grant it per deployment on that basis.
Set resources where a pod's appetite matters. No integration pod carries
requests or limits by default, and this is the one workload whose purpose is
running other programs. GKE Autopilot requires them.
Dev runs
The editor's Run button, executed as a pod the orchestrator creates rather than as a child process of whichever platform replica answered the request. The platform runs several replicas with no session affinity, so a run living in one replica's memory would be invisible to the others. On by default; disabling it makes Run unavailable. See Dev runs for the feature.
| Key | Default | Purpose |
|---|---|---|
orchestrator.devRuns.enabled | true | Renders the two images, the sidecar port, the idle timeout and the key Secret into the orchestrator's environment. Off, the orchestrator reports dev runs unavailable and the editor's Run says so. |
orchestrator.devRuns.hashSecret | "" | Required. HMAC key deriving each run's identity and its public hostname from (user, integration). |
orchestrator.devRuns.existingSecret / .existingSecretKey | "" / dev-run-hash-secret | That key from a Secret you already hold. The requirement is on the pair, not on hashSecret alone. |
orchestrator.devRuns.idleTimeout | 60m | How long a run survives untouched before the orchestrator reaps it. The only bound on how many pods an editing session leaves behind; there is no per-user cap. |
orchestrator.devRuns.sidecarPort | 8099 | Port the run's sidecar serves its reload/status API on, inside the pod. Never reached from outside it. |
The key must never change. A run's public hostname is derived from it, so
rotating it re-labels every exposed dev run and a webhook registered against the
old hostname silently stops being delivered. The chart requires the value
instead of generating one, because a generated key would be regenerated by any
render that cannot read the cluster (helm template, --dry-run, a GitOps
pipeline). Hold it wherever you hold kv.encryptionKey.
A Secret rather than a values file is the better answer for every credential the
chart takes. existingSecret is the same shape everywhere:
postgres:
auth:
existingSecret: octo-db-password # key: postgres-password
kv:
existingSecret: octo-kv-key # key: kv-encryption-key
auth:
existingSecret: octo-auth-creds # keys: oidc-client-secret, auth-secret
embeddings:
existingSecret: octo-embeddings-key # key: apiKey
orchestrator:
devRuns:
existingSecret: octo-devrun-key # key: dev-run-hash-secretDo not name your Secret after one the chart creates. Those are
{release}-postgres, {release}-auth, {release}-kv, {release}-devruns and
{release}-embeddings. {release}-postgres collides immediately, since the
chart creates it whatever the password's source and Helm refuses to install
over a Secret it does not own. The other four collide when you migrate: the
upgrade that stops rendering a Secret the previous revision owned deletes the
Secret your new revision points at. Neither shows up in helm template.
With references, no key material lands in the release history, and the chart
creates no Secret of its own for the auth, KV, dev-run or embeddings credentials.
{release}-postgres is still created because it carries the username and
database name; with postgres.auth.existingSecret set it holds those two and no
password.
--set-file and set_sensitive do not help. They change how a value is
supplied and how it is printed, not where it ends up: Helm writes every value it
is given into the release Secret in the cluster, and keeps it for every retained
revision. A reference is the only form that does not.
Clear the inline value when you switch: setting both is refused, since leaving
encryptionKey, hashSecret, apiKey or the two auth values beside an
existingSecret would keep the copy the switch was made to avoid. Revisions
written before the change still hold what they were given and helm history
keeps them, so migrate before anything has been encrypted or deployed.
A dev-run hostname is stable but not private. It is an unguessable hash, not
an access control. With orchestrator.clusterIssuer set, each one also gets a
per-host certificate and therefore an entry in public CT logs, so prefer
wildcardTLS (one *.{baseDomain} certificate) when dev runs are on. Without
orchestrator.baseDomain a run still works, reachable only in-cluster.
Data retention
A CronJob that asks the observability service, once a night, to delete stored logs and traces older than the site's retention policy. These values decide whether the job exists and when it runs; how long to keep things is a policy in the database, edited in the admin section. See Data retention.
| Key | Default | Purpose |
|---|---|---|
retention.enabled | true | Install the CronJob. The policy defaults to keeping everything, so a fresh install sweeps nightly and deletes nothing until somebody sets a window. Turning it off removes the schedule, not the capability: the endpoint stays served, so an on-demand sweep still works. |
retention.schedule | 0 3 * * * | When it runs, in the cluster's timezone. The first sweep after a policy is set can be the whole of both tables, and it competes with ingest while it runs. |
retention.timeoutSeconds | 900 | How long curl waits. It matches the service's own limit on a sweep, so neither side abandons a purge that is making progress. |
retention.concurrencyPolicy | Forbid | Skip a run while the previous one is still going, rather than stacking sweeps that would only refuse each other with a 409. |
retention.backoffLimit | 2 | Retries within one scheduled run. Each batch commits on its own, so a failed sweep is less deleted rather than inconsistent, and tomorrow's run picks up where it stopped. |
retention.image.* | curlimages/curl | The job is one HTTP POST and the octo images are distroless (no shell, no client). Third-party, so it takes its registry from retention.image.registry alone (empty means Docker Hub) rather than from the shared image.registry, as postgres and nats do. |
Scaling
| Key | Default | Purpose |
|---|---|---|
platform.replicas / orchestrator.replicas / observability.replicas | 1 | Replica counts. The observability service scales safely: its consumers join NATS queue groups, so replicas compete for messages rather than each storing every one. A retention sweep takes a Postgres advisory lock, so extra replicas cannot sweep concurrently either. |
nats.enabled | true | Deploy the NATS broker; when off, live event streams fall back to polling. See Event bus. |
redis.enabled | true | Deploy Redis. Unlike NATS this cannot simply be turned off; see Redis. |
externalRedis.url | (none) | A redis:// or rediss:// URL, with no credential in it, for a Redis this chart does not run. Rendered as a literal env value. |
externalRedis.existingSecret | (none) | A Secret you created holding the whole URL, for a server that wants a password. Bound by secretKeyRef, so the credential never enters a workload template. Wins over url. |
externalRedis.existingSecretKey | redis-url | The key within that Secret. |
redis.maxmemory | 256mb | The ceiling the server holds itself to. Keep it below any container memory limit, or the kernel gets there before Redis does and the eviction policy never runs. |
redis.maxmemoryPolicy | allkeys-lru | What it does on reaching the ceiling. Evicting the coldest keys costs the folds they held; growing without bound costs the pod. |
runtime.servicesModule | k8s | Runtime-services backend injected into integration pods: k8s (Lease leader election + orchestrator KV), standalone, or api (delegated to a server you implement; set runtime.repository to the octo-api image and inject OCTO_PLATFORM_API_URL too). Cluster deploys want k8s. |
Redis
Redis holds state shared between replicas of a component. Today that is one
thing: the observability service folds the trace records a streaming block emits
(one block.pre-invoke and one block.post-invoke per streamed token) into a
single row. Its consumers are a NATS queue group, so one trace's records are
spread across replicas.
It is not optional the way NATS is. nats.enabled=false only degrades live
event streams to polling, but the aggregator refuses to start without a Redis, so
an install with neither the bundled server nor externalRedis.url fails while
rendering, with a message naming both values.
Nothing here is persisted: no volume, and the server runs with --save "" and
--appendonly no. Losing it costs at most the trace records in the folds that
were open. To use a managed Redis instead:
redis:
enabled: false
externalRedis:
url: redis://cache.internal:6379The whole connection is that one string. If it contains a password, put it in
a Secret instead. externalRedis.url is rendered into the workload as a
literal environment value, readable by anyone who can read workloads:
kubectl create secret generic octo-redis \
--from-literal=redis-url='rediss://default:pw@cache.example:6380'redis:
enabled: false
externalRedis:
existingSecret: octo-redis # wins over `url` when both are set
existingSecretKey: redis-url # the defaultThe chart will not create that Secret for you. Whether the cluster's Redis is reachable is shown on the Platform services page, alongside Postgres, NATS and the Kubernetes API.
What the chart does not create
Four things a production chart often ships are deliberately absent, each with a workable answer in the meantime:
| Not shipped | Why | If you need it |
|---|---|---|
| PodDisruptionBudget | A PDB on a single-replica Deployment either permits the eviction it was meant to prevent or blocks node drains outright. It becomes meaningful once you raise replicas, differently per component. | Apply your own alongside the release; the components carry standard app.kubernetes.io/component labels to select on. |
| HorizontalPodAutoscaler | Throughput lives in the integration pods the orchestrator creates, not in the chart's own workloads, so an HPA here would scale the wrong thing. | Apply your own; platform and observability scale horizontally without coordination, the orchestrator is safe to scale but gains little. |
| NetworkPolicy | Correct policies depend on what else runs in the namespace and which CNI enforces them, and a policy that is silently unenforced reads as protection that is not there. | Write them against the component labels. The orchestrator needs the API server; every service needs Postgres and NATS. |
| Topology spread / anti-affinity by default | The knobs exist (defaults.topologySpreadConstraints, defaults.affinity) and the GKE Standard profile uses them; defaulting them would make single-node clusters unschedulable. | Set them per environment; see values-gke-standard.yaml. |
Image and chart compatibility
A chart at version X expects the images of release X, which is what it installs
by default via the appVersion fallback. This matters because the images run
as non-root, and the chart's cloud profiles pin the matching UIDs:
| Image | Runs as |
|---|---|
octo-platform-paas | 1000 (node) |
octo-orchestrator-paas, octo-observability-paas, octo-embeddings-paas, octo-runtime-paas | 65532 (distroless nonroot) |
octo-schema-paas, postgres | 70 (postgres) |
redis | 999 (the image's own user) |
nats | any UID (the profiles pin 1000) |
Pinning image.tag to an older release while using a newer chart is untested,
and the security contexts in the profiles are the most likely thing to break;
relax defaults.securityContext accordingly.
Cluster prerequisites
The chart deploys workloads, not cluster infrastructure. Depending on the values you choose, the cluster must already provide:
- an ingress controller matching
ingress.className; none is installed here - cert-manager and a ClusterIssuer, for
tls.mode: cert-manager(and forwildcardTLS, which needs a DNS-01 issuer) - a StorageClass for the Postgres PVC, unless
postgres.storage.hostPathis set orpostgres.enabledisfalse
Upgrade
helm upgrade octo oci://ghcr.io/juancavallotti/charts/octo \--version 0.11.7 --namespace octo --reuse-valuesBecause image.tag follows appVersion, a chart-version bump moves the images
with it. What happens:
- changed image references rewrite the pod templates, so the platform, orchestrator and observability Deployments roll automatically
- the schema hook Job re-runs and applies
sql/schema.sql; the schema is idempotent (IF NOT EXISTS/ON CONFLICT), so this is safe every time - Postgres and its PVC are untouched, so your data survives the upgrade
- integrations already deployed keep running on the runtime image they were
deployed with; redeploy them from the editor to pick up a new
octo-runtime
In the GCP reference deployment, Terraform owns the Helm release. Do not run
helm upgrade against it by hand; apply the release root (or
task deploy TAG=…) instead.
Validating changes
task helm:test lints and renders every profile and pipes the output through
kubeconform against the real Kubernetes schemas. CI runs it on every push, so a
template change that breaks a target you cannot reach locally still fails on the
pull request.