Octov0.11.7
Runtime

Observability

Health probes and Prometheus metrics on the runtime's admin port.

By default your runtime serves an admin port, separate from any http connector, carrying liveness and readiness probes and, when you ask for them, Prometheus metrics. Turn it off with --observability=false. This page covers what it exposes, what the numbers mean, and what they cost.

The admin port

octo run --config app.yaml
# ... the admin port is already listening on :39999

curl localhost:39999/healthz   # ok
curl localhost:39999/readyz    # ready
Endpoint
GET /healthzLiveness: 200 ok, always, while the process is running.
GET /readyzReadiness: 200 ready, or 503 naming the state.
GET /metricsPrometheus exposition. Served only with --metrics; 404 otherwise.
GET /An index of what is mounted.

It is a port of its own rather than a path on your http connector so that a probe never collides with one of your routes, never inherits a flow's CORS rules or timeouts, and answers for a runtime with no HTTP source at all. A cron- or queue-driven integration is exactly as probeable as a web one.

FlagEnvironmentDefault
--observabilityOCTO_OBSERVABILITYtrue
--observability-addrOCTO_OBSERVABILITY_ADDR:39999
--metricsOCTO_METRICSfalse
--metrics-blocksOCTO_METRICS_BLOCKS(none)

On the platform, enabling Pod Stats sets OCTO_METRICS=true on every deployed integration, because the sidecar it injects has nothing to scrape otherwise. It is set last in the container's environment, so it wins over an OCTO_METRICS binding in the deployment's own variables.

Each flag's default is its environment variable, so passing the flag always wins. These variables are read from the process environment only, not from a .env file: they are resolved when octo run parses its flags, before any config is loaded. --observability=false binds nothing at all.

When the port is already taken

In production a runtime owns its pod and :39999 is free. On a development machine the second octo run on one host finds the first one's admin port taken. That is not a startup failure: the runtime logs the bind error and carries on serving your flows with nothing on the admin port.

Give each run an address of its own to have both serve probes:

octo run --config app.yaml &
octo run --config other.yaml --observability-addr :39998

The editor's RUN does this for you: each run it spawns gets its own admin port, so several integrations can be running and monitored side by side.

Liveness and readiness are different questions

Liveness asks whether the process is wedged. /healthz answers 200 unconditionally: the handler running at all is the answer. There is no gate behind it, because a process too stuck to answer is what a liveness probe already detects by timing out.

Readiness asks whether the runtime is serving. It has a real answer, and it changes:

State/readyzWhen
starting503The process is up; connectors and flows are not.
ready200Every connector and flow started.
reloading503Under --watch: a config change is being rebuilt, or the config will not load.
draining503A shutdown signal arrived; flows have not finished.
stopped503The runtime has stopped.

The state is in the response body, because "why is this pod not ready" is the next question and a bare 503 sends you to the logs to find out.

Two of these matter more than they look. The admin server binds before connectors start, so a probe during starting gets a real "not yet" instead of a refused connection. And draining is set the moment a shutdown signal arrives, before flows drain, so a load balancer sees the pod leave the pool while workers are still finishing what is in flight.

ready means every connector and every flow started, which for an HTTP-backed flow includes its listener being bound. Readiness is therefore a true answer for cron and queue runtimes, not an HTTP-shaped approximation.

On the platform you get this for free: the orchestrator points each deployment's livenessProbe and readinessProbe at this port, on every deployment, networked or not.

Metrics

octo run --config app.yaml --metrics
curl -s localhost:39999/metrics

Per flow

MetricTypeLabels
octo_flow_messages_totalcounterflow, outcome (completed/dropped/failed)
octo_flow_errors_totalcounterflow, block
octo_flow_duration_secondshistogramflow, outcome
octo_flow_in_flightgaugeflow

octo_flow_in_flight is the one to put on a dashboard first. It is the number of messages a flow is processing right now, and the only signal that distinguishes "quiet" from "stuck". A flow whose workers are all blocked on a slow database shows in-flight pinned at its worker count (8 or more, see worker pools) while throughput and CPU both sit flat.

octo_flow_errors_total counts unhandled failures only. A message that failed and whose flow-level error path recovered it is reported as completed, so it is not counted here. The metric is therefore an error rate, and a flow that handles its errors reads as healthy.

The block label names the block the failure came from, so you get "the charge block is failing" rather than "this integration is failing". It is empty for a failure that did not originate in a block.

A flow that has taken no traffic still exports its series at zero, so a silent flow is visible as 0 rather than missing. Reloading the config under --watch only ever adds series; counters are never reset, because a reset makes rate() invent a spike.

Per block

Off by default. Name the blocks you want, using the same address grammar --spies accepts:

octo run --config app.yaml --metrics \
  --metrics-blocks 'orders.charge,orders.fanout[audit].log-it'
MetricTypeLabels
octo_block_invocations_totalcounterflow, path, type, outcome (ok/dropped/error)
octo_block_duration_secondshistogramflow, path, type
octo_block_listener_panics_totalcounter(none)

This is the difference between "this integration is slow" and "the sql block in this integration is slow". * watches every block, which is the fastest way to find where the time goes on a small config and a lot of series on a large one.

A watched block is measured on the flow's own goroutine, in the middle of the block it is timing, so the measurement is on the message's critical path. It is two atomic updates and no allocation, small against any block that does real work but not free, which is why this is a separate flag from --metrics. Blocks you did not name cost one comparison each; '*' puts the cost on every block in every flow and is a debugging tool, not a steady-state setting.

Two things to know before you trust the numbers. The counters are exact: every watched invocation is recorded inline, so nothing is sampled and nothing is dropped however hard the runtime is being pushed. The one way to lose a measurement is a listener that panics, which is contained, logged with its stack, and counted in octo_block_listener_panics_total (see block events); that counter should be zero.

The second is that a composite's duration includes its children's, and the children emit their own events, so summing octo_block_duration_seconds across paths double-counts. The same applies across flows: a flow reached by flow-ref times itself, and its duration nests inside its caller's.

Process and build

Metric
octo_ready1 when serving, 0 otherwise: the same signal /readyz answers, read at scrape time so the two cannot disagree.
octo_flows, octo_connectorsThe size of the running configuration.
octo_build_infoAlways 1; the labels carry version, build_date and services_module.
go_*Goroutines, heap, allocation and GC behaviour.
process_*Resident memory, CPU seconds, open file descriptors.

The go_* and process_* families come from the standard Prometheus collectors, and answer most "why is this pod using that much memory" questions.

Cardinality and cost

Every label is filled from your config (a flow's name, a block's address, a fixed outcome) and none from message data. Cardinality is therefore bounded by what you wrote, not by your traffic: no label can be a customer id, a URL path, or an error string.

The costs, in the order they matter. Probes are free, a listener and two handlers. --metrics is cheap: the flow-event bus is already publishing whether or not anything listens, so the added work is a counter increment, a gauge step and one histogram observation on the flow's own goroutine. --metrics-blocks is the one with a real cost, per the warning above. Scraping walks the registry, and is worth watching only if you watch a very large number of blocks.

Prometheus configuration

prometheus.yml (excerpt)
scrape_configs:
  - job_name: octo
    static_configs:
      - targets: ['octo-host:39999']

Useful starting queries:

# Throughput per flow
sum(rate(octo_flow_messages_total[5m])) by (flow)

# Error rate per flow, as a fraction
sum(rate(octo_flow_errors_total[5m])) by (flow)
  / sum(rate(octo_flow_messages_total[5m])) by (flow)

# Which block is failing
topk(5, sum(rate(octo_flow_errors_total[5m])) by (flow, block))

# p95 latency per flow
histogram_quantile(0.95,
  sum(rate(octo_flow_duration_seconds_bucket[5m])) by (le, flow))

# Saturation: every worker busy, throughput flat
octo_flow_in_flight

See also

  • Monitoring, for logs, the log block, and the event seams these metrics are built on.
  • CLI reference, for the flags and environment variables.
  • Processing pipeline, for what a flow event is and when each one fires.

On this page