Observability API Reference
The observability service describes its own API: an OpenAPI 3.1 document, filters for reading part of it, and a compact operation index.
The observability service serves a description of its own API, the same way the orchestrator does. Everything the /platform/logs, /platform/traces and /platform/metrics views read goes through that API, and so do the retention settings and the storage health tab, so the description is the reference for querying stored telemetry yourself.
curl -sS "$OBSERVABILITY_URL/openapi.json"It is OpenAPI 3.1, generated from annotations on the handlers at build time, committed, and embedded in the binary. It is never generated at runtime, so what a running service serves is exactly what was reviewed with the code that serves it.
The service began as a log aggregator and grew: traces, pod stats, the retention policy and the storage report all landed here because each is records shipped in and history queried out, and one service holds the stores they live in. It is observability/ in the repository, octo-observability in the cluster, and OBSERVABILITY_URL to everything that reaches it.
The operation index
/openapi/operations is a compact list of every route: method, path, summary and tags.
curl -sS "$OBSERVABILITY_URL/openapi/operations"[
{
"method": "GET",
"path": "/logs",
"summary": "List log events",
"tags": ["logs"]
}
]Sorted by path, then method. Read this first to find the tag worth asking about, then fetch that tag's detail, which is how the platform agent uses it. The tags are logs, traces, stats, alerts, retention, storage, meta and openapi.
Reading part of the document
Two filters, mutually exclusive; supplying both is a 400:
| Query | Keeps |
|---|---|
?tag=traces | Operations carrying that tag |
?path=/traces | Routes whose path starts with that prefix |
curl -sS "$OBSERVABILITY_URL/openapi.json?tag=logs"A filtered document is a valid OpenAPI document, not an excerpt: schemas are pruned to the ones the surviving operations can reach, following $ref transitively, so the result still resolves. An unknown tag yields a document with no paths rather than an error.
What the API offers
Every query route is a GET, and so is the storage report. The three retention routes are the exception, and one of them deletes; so are the alerting routes.
Paging is keyset-based. /logs and /traces return at most limit rows (default 200, clamped to 1000). A page that came back full carries next_before; send it as before= for the page after it. Treat it as opaque: it names a (timestamp, id) pair, not a timestamp, because rows routinely share a timestamp.
Trace payloads are optional. /traces/{traceId} includes every record's captured body by default, and those payloads are the whole weight of a trace. Pass bodies=0 when you are only drawing a waterfall, then fetch the one record you need from /traces/{traceId}/records/{id}.
Pod stats
The pod stats sidecar writes each production pod's own metrics to Redis. These three routes read them back, and they are the only part of this API not backed by Postgres, which is why they answer even when the log store is unreachable.
| Route | Purpose |
|---|---|
GET /stats/{deploymentId}/pods | Which pods have reported, and whether they still are |
GET /stats/{deploymentId}/metrics | The catalogue: every series name and label set, no samples read |
GET /stats/{deploymentId}/series | The data |
Start at /pods. It names the pods the other two accept, and reports liveRows and rollupRows separately: zero live rows beside a full history is the ordinary state of a pod that stopped a few hours ago, because the live tier is kept for only twice the rollup interval while the pod stays in the index for the whole retention window.
metric is required on /series, and repeatable. Samples are stored positionally against a per-pod dictionary, so a query with no name filter would read every series of every pod, around ninety-five of them, once a second. /metrics is how you find exact names to put in it. Names match exactly; label=key=value narrows within them and repeats are ANDed.
curl -sS "$OBSERVABILITY_URL/stats/$DEPLOYMENT_ID/series?metric=process_resident_memory_bytes&from=2026-09-05T11:00:00Z"A counter's value is growth, on both tiers. A Prometheus counter is cumulative, so what is charted is how much it grew over the interval ending at that point. The history tier stores that already; the live tier is differenced to match. A counter that reads lower than the point before is treated as a process restart rather than a negative delta. Pass counters=absolute for the raw cumulative readings, or ask for the last stat on the rollup tier.
tier=auto is the default and resolves to rollup whenever the window reaches further back than the live tier goes. That reach is liveDepth × sampleInterval, one hour at the defaults; a window older than it cannot be answered from live rows however many are read. The resolved tier and its step come back in the response.
Points are columnar, and times are unix milliseconds. The window in the envelope stays RFC3339, and so do the from and to parameters. Ask for only the numbers you need with stats=value,max; the extra columns are rollup-only and omitted when not requested.
A null reading is a gap, not a zero: a series the dictionary knows that the scrape did not report. On the rollup tier, ends is carried beside times; rows are not contiguous, so a bucket's end not meeting the next one's start is where scraping stopped.
A deployment with no stats answers 200 with no items rather than 404. This service holds no deployment registry, so it cannot tell a deployment that never existed from one whose sidecar is switched off or whose stats have expired. For the same reason a pod that cannot be read is reported in warnings and skipped, never a failed request.
Data retention
Stored telemetry grows without bound until you say otherwise. The policy lives here rather than on the orchestrator because this service owns the tables a sweep deletes from. Set it at /platform/admin/retention unless you are automating it; these are the routes that page calls.
| Route | Purpose |
|---|---|
GET /settings/retention | The three windows, in days, and when they were last changed |
PUT /settings/retention | Set them |
POST /retention/run | Enforce the stored policy now |
curl -sS -X PUT "$OBSERVABILITY_URL/settings/retention" \
-H 'Content-Type: application/json' \
-d '{"logs_days":30,"traces_days":7,"alerts_days":14}'Zero means keep forever on every axis, and the upper bound is 3650 days. Zero is also the default for logs and traces. alerts_days defaults to 14: a watch records an evaluation every time it runs whether or not anything happened, so its history grows on a schedule rather than with traffic. Setting it to zero explicitly still means forever.
A save replaces the whole policy, so all three fields are required and a body naming fewer is a 400. Zero is a value here, not an absence, so a request mentioning only logs_days would otherwise read as "and switch the other two off".
A save deletes nothing. POST /retention/run is what deletes, and it reports what went:
{
"logs_deleted": 120412,
"traces_deleted": 44980,
"trace_summaries_deleted": 2211,
"alert_evaluations_deleted": 201600,
"alert_incidents_deleted": 18,
"logs_cutoff": "2026-07-12T03:00:00Z",
"traces_cutoff": null,
"alerts_cutoff": "2026-07-28T03:00:00Z",
"duration_ms": 8140
}A null cutoff tells an axis set to keep everything apart from one where nothing had expired; both report zero deleted. A policy that keeps everything is a 200 with zeros rather than an error.
Only one sweep runs at a time, held by a Postgres advisory lock, so the nightly job and somebody pressing Run now cannot overlap. The loser gets a 409.
Deletes are issued in batches and each commits on its own, so an interrupted sweep has simply deleted less and the next run continues from there. A trace is deleted whole, and stored model prices are never swept; see Traces. An open alert incident is never swept either, whatever its age: it is current state rather than history, and deleting one would leave a firing watch pointing at an episode that no longer exists.
Running it nightly
The chart installs a CronJob that calls POST /retention/run at 03:00, controlled by retention.enabled (on by default) and retention.schedule. Because the two evidence windows default to keeping everything, that job only sweeps the alerting history until you set one; turning retention on for logs or traces is a change to the policy, not to the release.
Storage report
GET /settings/storage reports how full the two stores are: Redis memory against its ceiling, the hit rate, evictions and expiries; and for Postgres this service's connection-pool usage, the database size and the size of the object store's kv_store table. It is what the Storage health tab beside the object browser renders.
It lives here because this service holds both stores and is the heaviest writer to one of them: every log and trace record lands through the pool it reports, so a pool with no connection to spare is the shape of telemetry backing up. Either half is null with a reason beside it when the installation does not have that store or cannot reach it, and the two reasons are different because an installation with no Redis is a supported one.
curl -sS "$OBSERVABILITY_URL/settings/storage"The dependency health report stays on the orchestrator, because whether the orchestrator can reach what it needs is a question only the orchestrator can answer.
Keeping it honest
The spec is regenerated in CI and the build fails if the committed copy differs, the same way it fails on unformatted code.
A separate test compares the routes registered in the source against the routes described in the spec, in both directions. Regenerating cannot catch a route that was added without annotations, so a new route with no documentation fails that test by name.
Authentication
Every route needs a bearer token minted by this installation's iam service, except /healthz and the two description endpoints. The service fetches iam's published keys through discovery and verifies the signature, the issuer, the audience and both time claims itself.
What a route requires depends on the roles the token carries:
| Routes | Read | Write |
|---|---|---|
/logs, /traces, /stats | any people role | — |
/alerts/incidents/{id}/ack | any people role | any people role |
/alerts/watches, /alerts/preview | any people role | platform:admin |
/settings/retention, /settings/storage, /retention/run | platform:admin | platform:admin |
"Any people role" is platform:admin, platform:operator, platform:developer or platform:monitor — the four a person can be granted. platform:runtime, which is what a deployment's own token carries and no person ever holds, is in none of these rules, so such a token reads nothing here.
Anything matching no rule is platform:admin, so a route added without a thought about access fails closed.
A refusal is 401 when the token is missing or bad and 403 when it is good and does not carry enough. A 503 means the service could not reach iam to decide — deliberately not 403, because a caller told they are forbidden goes and changes permissions that were never the problem.
Reading stored traces means reading the request bodies of everything traced on this installation, not just your own deployments. The policy above admits every role to that read, because a monitor who cannot see what they are monitoring has no job. Grant platform:monitor on those terms.
An installation with IAM_URL unset authorizes nothing and serves every caller, which is what a local run is. In the chart it is always set.
The browser never reaches any of this directly. The platform's server actions do, presenting the signed-in person's own token, so what a person can read here is what they can read in the UI.