Octov0.11.7
Extending

The platform API

Run Octo on your own platform by implementing one HTTP contract.

Octo's platform capabilities (the object store, secrets, queues, topics, leases, leader election, agent memory, traces and logs) normally come from the runtime itself or from Octo's own PaaS. The platform API is the third option: you implement one HTTP contract, and every one of those capabilities is delegated to you.

It is a consumer-defined interface. Octo defines the contract and is the only client; you implement it against whatever your platform already has. On Google Cloud that is usually Firestore, Secret Manager and Pub/Sub behind a small service of your own. On Kubernetes it can be a central service your pods call, or a sidecar on loopback.

You do not have to implement all of it. A server that answers the discovery endpoint and the three key/value operations is a useful implementation, and the rest degrades on the runtime's side without you writing a line for it. See What happens to what you skip.

When to reach for it

Reach for
One process, local disk, no infrastructurethe standalone provider (the default build)
Octo's own PaaS: the orchestrator, NATS, Kubernetes Leasesthe k8s provider
Cloud Run, Firebase, your own platform service, a sidecarthe platform API

The first two are described in Writing a runtime service. This page is the third.

Getting it running

Run the octo-api image, or any octo binary built with -tags api, and point it at your server:

docker run --rm \
  -e RUNTIME_SERVICES_MODULE=api \
  -e OCTO_PLATFORM_API_URL=https://platform.example.internal \
  -v "$PWD/integrations:/etc/octo/integrations" \
  juancavallotti/octo-api

A successful start says so, and says what it negotiated:

INFO api runtime services initialized url=https://platform.example.internal
     instance=octo-7b9d-xk21 deployment=orders
     implementation=acme-gcp-adapter implementationVersion=1.4.2
     features="[kv secrets leases leaderElection queues]"
INFO runtime services ready module=api

The features list is what your server said it implements, not what Octo wants. If a capability you expected is missing from that line, your discovery document is where to look.

The contract

The full contract is an OpenAPI 3.1 document. Start with the copy published here, octo-platform-api.yaml, while you are still deciding whether to implement it.

Once you are building against a particular runtime, take the copy out of that runtime instead. It is carried inside the octo-api image, so what you read is version-matched to the thing that will call your server:

docker run --rm juancavallotti/octo-api openapi > platform-api.yaml
docker run --rm juancavallotti/octo-api openapi --format json > platform-api.json

octo openapi exists only in the -tags api build, alongside the provider that speaks the contract; the standalone and cluster binaries carry neither. Read the YAML rather than the JSON: the comments are half the document, and the JSON rendering drops them.

Discovery

Everything begins at GET /v1/discovery. The runtime calls it once, at startup, and every other route is conditional on what it returns.

{
  "specVersion": "1.0",
  "implementation": { "name": "acme-gcp-adapter", "version": "1.4.2" },
  "features": {
    "kv": { "supported": true, "maxValueBytes": 1048576 },
    "secrets": { "supported": true, "encryptedAtRest": true },
    "queues": { "supported": true, "pollTimeoutSeconds": 20, "maxBatch": 8, "ackDeadlineSeconds": 60 },
    "leases": { "supported": true, "minTtlSeconds": 1, "maxTtlSeconds": 600 },
    "leaderElection": { "supported": true, "leaseTtlSeconds": 15, "renewIntervalSeconds": 5 }
  }
}

Durations are integers with the unit in the name rather than duration strings, so a "20s" parser is not needed in every implementation language.

Unknown fields are ignored, unknown features are ignored, and a feature you leave out is a feature you have not implemented. A newer server talking to an older runtime is fine, and a capability added to this contract later reaches you as something you simply do not mention.

Discovery is fetched once and is not re-polled. What does exist is the 501 latch: any route answering 501 Not Implemented turns that capability off for the life of the process, logged once. That covers a runtime rolled forward onto routes your server has not grown yet.

If the runtime cannot reach discovery at all it retries with backoff for 30 seconds before giving up. See the sidecar recipe for why that window matters.

What happens to what you skip

An unsupported capability either degrades (reads come back empty and writes return a named error) or refuses every call, reads included. Most features let you choose with "unsupported": "noop" or "unsupported": "error"; the defaults, and the three that have no choice, are below.

FeatureDefault when unsupportedWhat the runtime does
kvdegradeReads miss; writes and deletes return kv: no store configured.
secretsdegradeSecrets are a view over the *_secrets namespaces, so a working kv keeps them working whatever this says. Set "unsupported": "error" to refuse them instead, which is how you say "I store values but not secrets".
resourcesdegradeEvery .env file and template reports missing.
leasesrefuseEvery claim fails with a named error.
leaderElectionrefuseEvery campaign fails with a named error.
queues / topicsdegradePublishes fail loudly; subscriptions are inert.
agentMemorydegrade, alwaysAgents keep no memory and take their pre-memory path.
traces / logsdegrade, alwaysNothing is shipped.

The last two rows say always because those three have no refusing mode and "unsupported": "error" on them is ignored. The engine checks whether memory is enabled before it would ever reach a write, and traces and logs are best-effort by construction.

Leases and leader election refuse by default, and everything else degrades. Degrading those two would mean granting them: a no-op lease hands out every claim and a no-op election makes every replica the leader. That is correct for a single process and exactly wrong behind this module, where two Cloud Run instances would each be told they hold the claim and both would run the work the claim exists to run once.

If you genuinely run one instance, say so with "unsupported": "noop" and the single-process semantics are right again.

The five things implementations get wrong

The contract documents every route. These five are what people actually get wrong, and each fails in a way that is hard to trace back:

404 means the addressed thing is absent, never "no such route". A missing key, a released lease, a conversation that does not exist yet. A route you have not implemented answers 501, and the runtime turns that capability off rather than failing every call against it forever.

204 on a receive is the long poll expiring, not an error. It is the normal state of an idle subject. Hold the request open for the waitSeconds you were sent, and answer 204 only when that window passes with nothing to deliver. A server that answers immediately turns every subscription into a busy loop; the runtime defends itself by waiting out the rest of the window.

X-Object-Version: 0 on a write means create. It must fail with 409 if the object already exists; a positive value must equal the stored version, or 409 again. Getting this wrong does not fail loudly. It silently loses concurrent updates, which is the entire thing the check exists to prevent.

Keys and names arrive as query parameters, not path segments. They may contain a slash, and %2F inside a path segment is normalized by nginx, by Cloud Run's front end and by several frameworks, quietly merging two distinct keys into one. Do not move them into the path.

userId on the agent-memory thread routes matters. Record who a conversation is with on the first write that names one. Ignoring it stores a complete, correct history attributed to nobody, and the person's own view of their conversations then shows as empty.

There is no pub/sub here

Queues and topics in this contract are pull: the runtime long-polls you, acknowledges what it handled, and rejects what failed. Nothing is ever pushed to the runtime.

That fits what Octo's queues are for: distributing work between the replicas of one deployment, smoothing a burst, running a slow step off the request path. The acknowledgement makes delivery at-least-once, which is stronger than the other two providers offer.

It is not an event bus. If something outside needs to trigger a flow, do not model it here. Give the flow an HTTP source and have your event source call it:

orders.yaml (excerpt)
sources:
  - type: http
    settings:
      path: /events/order-created
      method: POST

Then point a Pub/Sub push subscription, an EventArc trigger, a webhook, or anything else that speaks HTTP at that path. That is the supported pattern for external ingress, it needs nothing from this contract, and it gives you the retry and delivery semantics of whatever bus you already run.

Pull for work between your own replicas, HTTP for anything arriving from outside. Topics over this contract exist for internal fan-out, not for subscribing to your platform's event stream.

Configuration

Every setting the module reads. Only the first is required.

VariableDefaultWhat it does
OCTO_PLATFORM_API_URLrequiredYour server's base URL.
OCTO_PLATFORM_API_TOKENemptyBearer token, sent on every request.
OCTO_PLATFORM_API_TOKEN_FILEemptyRead the token from a file instead, re-read when it changes. Wins over _TOKEN.
OCTO_PLATFORM_API_HEADERSemptyExtra static headers as Name: value, separated by newlines or commas.
OCTO_PLATFORM_API_CA_FILEemptyA custom CA to trust.
OCTO_PLATFORM_API_CLIENT_CERT_FILEemptyClient certificate, for mTLS.
OCTO_PLATFORM_API_CLIENT_KEY_FILEemptyClient key, for mTLS.
OCTO_PLATFORM_API_TIMEOUT10sDeadline for an ordinary request.
OCTO_PLATFORM_API_LONG_TIMEOUT60sCeiling for long polls and memory search.
OCTO_PLATFORM_API_STARTUPrequirerequire fails startup if discovery never answers; degrade starts with every capability unavailable.
OCTO_PLATFORM_API_DISCOVERY_BUDGET30sHow long to retry discovery before applying the startup policy.
OCTO_DEPLOYMENT_IDemptySent as X-Octo-Deployment on every request, and used as the queue consumer group so replicas of one deployment compete. What else you scope by it is yours to decide. Reused from the cluster provider, not invented here.
OCTO_INSTANCE_IDPOD_NAME, else hostname and pidThis replica's identity, for leases and leader election.

OCTO_PLATFORM_API_HEADERS is the escape hatch for an API key header, a tenant id, or a service-mesh header, none of which has a variable of its own.

OCTO_INSTANCE_ID falls back to POD_NAME so a Kubernetes Deployment that already projects the downward API gets working leases with no extra wiring. On Cloud Run the hostname is per-instance, and the pid disambiguates two runtimes sharing one.

Checking your implementation

The octo-api image can check a server against the contract and tell you, rule by rule, what it satisfies:

docker run --rm juancavallotti/octo-api \
  verify-platform-api https://platform.example.internal
Platform API at https://platform.example.internal
  implementation: acme-gcp-adapter 1.4.2
  specVersion:    1.0 (this runtime speaks 1.0)
  declared:       kv secrets queues

  ok    kv             reading an absent key is a miss, not an error
  ok    kv             version 0 creates
  FAIL  kv             version 0 over an existing key conflicts
                       0 means create; writing it over an existing key must answer 409
  SKIP  leases         leases
  FAIL  queues         an empty poll waits before answering
                       the poll returned in 1ms of a declared 20s window. A server that
                       answers immediately makes the runtime poll as fast as the network allows

  8 passed, 2 failed, 1 skipped

Like octo openapi, it is in the -tags api build only, and it drives the same client the runtime does, so what it checks is what will actually happen. SKIP is not a failure: a feature you did not declare is one nothing is sent to. It exits non-zero on any FAIL, so it belongs in the pipeline that ships your server, and --json gives you the report as data.

It writes, under an octo-verify/ prefix it names before it starts. Point it at staging.

A minimal implementation

Discovery and key/value storage, enough to run flows that use the object store and secrets, with everything else degrading. Around a hundred lines in either language.

import express from 'express';

const app = express();
const store = new Map<string, { value: Buffer; version: number }>();
const keyOf = (ns: string, key: string) => `${ns} ${key}`;

app.get('/v1/discovery', (_req, res) => {
  res.json({
    specVersion: '1.0',
    implementation: { name: 'example', version: '0.1.0' },
    features: {
      kv: { supported: true },
      secrets: { supported: true, encryptedAtRest: true },
      // Explicit, not omitted: this deployment runs one instance, so granting
      // every claim is correct. Omitting these would make the runtime refuse.
      leases: { supported: false, unsupported: 'noop' },
      leaderElection: { supported: false, unsupported: 'noop' },
    },
  });
});

app.get('/v1/kv/:namespace/entry', (req, res) => {
  const row = store.get(keyOf(req.params.namespace, String(req.query.key)));
  if (!row) return res.status(404).end();          // a miss, not a failure
  res.set('X-Object-Version', String(row.version));
  res.type('application/octet-stream').send(row.value);
});

app.put('/v1/kv/:namespace/entry', express.raw({ type: '*/*' }), (req, res) => {
  const id = keyOf(req.params.namespace, String(req.query.key));
  const expected = Number(req.get('X-Object-Version') ?? 0);
  const row = store.get(id);
  // 0 means create, and must conflict if the key is already there.
  if ((expected === 0 && row) || (expected !== 0 && row?.version !== expected)) {
    return res.status(409).end();
  }
  const version = (row?.version ?? 0) + 1;
  store.set(id, { value: req.body, version });
  res.set('X-Object-Version', String(version)).status(200).end();
});

app.delete('/v1/kv/:namespace/entry', (req, res) => {
  const id = keyOf(req.params.namespace, String(req.query.key));
  const expected = Number(req.get('X-Object-Version') ?? 0);
  const row = store.get(id);
  if (row && expected !== 0 && row.version !== expected) return res.status(409).end();
  store.delete(id);                                 // deleting nothing is success
  res.status(204).end();
});

// Anything not implemented: 501, so the runtime turns that capability off
// rather than failing every call against it.
app.use((_req, res) => res.status(501).json({ error: { code: 'not_implemented' } }));

app.listen(8080);

Run octo verify-platform-api http://localhost:8080 against either and it should report the key/value checks passing and everything else skipped.

Deployment recipes

One image serves all three shapes, because they differ only in what OCTO_PLATFORM_API_URL points at.

Cloud Run

Two services: your platform API, and the runtime that calls it.

gcloud run deploy platform-api \
  --image gcr.io/my-project/platform-api \
  --no-allow-unauthenticated \
  --timeout 120

gcloud run deploy octo \
  --image juancavallotti/octo-api \
  --set-env-vars RUNTIME_SERVICES_MODULE=api,OCTO_PLATFORM_API_URL=https://platform-api-xxxx.run.app \
  --min-instances 1

Two flags matter more than they look:

--timeout on the platform API service must exceed your pollTimeoutSeconds. A long poll is a request held open, and Cloud Run's request timeout will cut it: a 20-second poll behind a 10-second timeout produces an error on every poll rather than the 204 the runtime is waiting for.

--min-instances 1 on the runtime matters if you use queues at all. A runtime scaled to zero is not polling, and Cloud Run scales on inbound requests, which a queue consumer has none of.

Authenticate with an ID token, or with a static header via OCTO_PLATFORM_API_HEADERS.

Kubernetes, as a central service

One platform-API Service; runtime pods reach it by name.

octo-runtime.yaml (excerpt)
env:
  - name: RUNTIME_SERVICES_MODULE
    value: api
  - name: OCTO_PLATFORM_API_URL
    value: http://platform-api.platform.svc.cluster.local
  - name: OCTO_INSTANCE_ID
    valueFrom:
      fieldRef:
        fieldPath: metadata.name

Leader election works across replicas because each pod sends a distinct identity. OCTO_INSTANCE_ID falls back to POD_NAME, so projecting either is enough.

Kubernetes, as a sidecar

Your platform API as a second container, reached over loopback:

octo-runtime.yaml (excerpt)
containers:
  - name: octo
    image: juancavallotti/octo-api
    env:
      - name: RUNTIME_SERVICES_MODULE
        value: api
      - name: OCTO_PLATFORM_API_URL
        value: http://127.0.0.1:8080
  - name: platform-api
    image: my-registry/platform-api
    ports:
      - containerPort: 8080

Container start order is not guaranteed, so the runtime will usually make its first discovery call before the sidecar is listening. It does not crash: it retries with backoff for OCTO_PLATFORM_API_DISCOVERY_BUDGET (30 seconds). That is why the sidecar shape needs no native-sidecar restartPolicy and no init container.

What differs from the other providers

standalonek8sapi
Queue deliveryat-most-onceat-most-onceat-least-once
Leader electionalways leaderKubernetes LeaseTTL and campaign polling
Listing agent conversationsyesnoyes, if you declare it

Queue delivery is at-least-once here, because a delivery is acknowledged after its handler returns and a failed one is handed back. A handler written against at-most-once still works, but one that is not idempotent may now see the same message twice.

Leadership is TTL-based. A leader that loses contact with your server stops asserting leadership one renew interval later, and that has to happen before its claim expires on your side, so a renew interval may be at most a third of the TTL, leaving room for two lost renewals. The runtime shortens one that is too close rather than trusting the document.

Listing conversations works here and not on k8s, because the cluster provider's pod knows only its deployment and listing is a tenancy question. Declare listThreads and readThread when your server should answer it.

Versioning

specVersion names the contract, and this runtime speaks 1.0. A mismatch is logged as a warning and nothing more.

Within a major version, Octo does not remove a route, change what a status code means, or make an optional field required. New capability arrives as a new feature block, which reaches you as something you do not mention in discovery, so nothing changes for an implementation that has not grown it yet.

See also

On this page