The Platform Agent
Install Dr. Octo, keep him updated, trace him, and edit him like any other integration.
Ask Dr. Octo about this installation and he answers: he reads and explains your integrations, writes and fixes flow definitions, and investigates deployments that will not start. Install him from Admin → Platform agent at /platform/admin/agent.
Dr. Octo is an ordinary integration. He is not a service, a sidecar, or a bundled image. He is a config.yaml with an ai-agent block in it, installed as a real integration and deployed like anything else. His logs, traces, scaling and version history are where every deployment's are, and the link on the status card opens him in the editor, where you can change the prompt, add or remove a tool, or rewrite a skill.
Before you install
Installing needs three things. The page tells you which one is missing instead of failing:
| Requirement | Why | If it is missing |
|---|---|---|
| Cluster access | He is deployed like any integration | Install is disabled; this orchestrator cannot deploy anything |
kv.encryptionKey | His provider key is read back out of storage to bind onto the deployment | Install is disabled, see Admin settings |
| A stored LLM key | There is no model for him to reason with | Install is disabled, with a link to the LLM provider section on the same page |
Install him
Press Install. That creates the integration, writes his skills as resources, publishes a version tag derived from the bundle this orchestrator ships, writes the provider key as a cluster secret, and deploys him with one replica and no external route. If the integration exists but nothing is running, the same button reads Deploy.
He is deployed with no platform access of his own, which looks wrong for an agent that drives the whole API and is not. When somebody is chatting, his tools spend that person's token, so what he may do is what they may do — and no grant on his deployment would change that.
His own token is what an unattended run falls back to: an alert waking the troubleshooter, where there is nobody to borrow from and never will be. He is installed holding both platform access grants for that reason — without them platform:runtime matches no rule on either API, so an alert-woken triage is refused even its reads, and an agent installed to investigate and repair could do neither. The grants change nothing about a conversation, where his tools spend your token rather than his; untick them on his deployment if you would rather unattended runs did nothing at all.
Installing is idempotent. An existing integration is reused and an existing tag is not recreated, so you never get a second agent, and retrying after a failed deploy picks up where it stopped. He is deployed internal-only: the platform proxies to the in-cluster address shown on the status card.
Dr. Octo has full read and write access to the orchestrator API, deploying included. The chat panel is therefore gated by the write roles — every role except platform:monitor, unless AUTH_WRITE_ROLES narrows it further. To narrow what he may do regardless of who asks, see Restrict what he can do.
Talk to him
Once he is running, a round button appears in the bottom-right corner of any signed-in page, but only when he is deployed and you hold the write roles. It opens a full-height drawer over the right-hand side of the page.
Type a question and the answer streams back a word at a time. The drawer shows:
- Markdown, so a flow definition arrives in a code block you can copy and a comparison as a table.
- His working, in order: what he thought, what he called, what he made of the result, then his reply. Reasoning collapses to a one-line summary once anything follows it; click the line to read it again.
- A running commentary above the message box: thinking, running a named tool, reading what came back, shortening the conversation.
- A context gauge in the header once a conversation is more than half way through his budget. He is about to start forgetting the beginning.
- Navigation. When the answer is a page, he takes you there and says why. The drawer refuses any destination outside the platform, whatever his definition has been changed to say.
- A question before he reads the web. The
web_searchchip opens with the objective and the queries he chose, theweb_extractchip with the URLs, and both wait for Allow or Deny. He holds the call while you decide, and is told after three minutes if nobody answered. See why those two.
Scrolling up stops the view following the answer; a Latest button appears while it is off, and returning to the bottom turns it back on.
Press Stop to end a run in progress, or close the panel. Closing aborts the request and the runtime stops the run when the connection goes. Stop is stronger: it sends a stop for the conversation, so it ends the run wherever it is running, including on a replica this browser never spoke to.
You can also type while he is working. A message sent mid-answer is handed to the run already in flight rather than starting a second one, and the answer goes to the stream you already have open. It shows as unread until he picks it up, which can be a whole model call later; then it moves into the conversation at that point, marked Read, and taken into account, with the reasoning and tool calls it caused underneath it. If the run ends first, or he runs out of steps before reaching it, the panel says so and tells you to ask again. A message from a second tab, or from a window that cannot tell a run is already going, still reaches that run; the panel says where the answer went.
Press New conversation to start fresh. That drops the thread and what he remembers of it, but the conversation is kept and listed by name in the panel. He knows which page you are on, so "this integration" or "this deployment" means what you think it means.
Conversations are scoped to you. The panel sends a thread id, but he keys his memory on the signed-in user and that thread, and the identity comes from your server session rather than from the browser. A thread id belonging to someone else names a conversation that does not exist.
Coming back to one
Every exchange is recorded as it happened, under a name he writes for the conversation on its first message. The history button in the drawer's header lists them; opening one replays it and continues the thread. What comes back is what was said, your messages and his answers; the reasoning and tool calls belonged to the run and went with it.
That record is not his memory. Memory is compacted (a long conversation has its early turns replaced by a summary), so it is the working set rather than the conversation you had. Both are readable side by side on the memory page.
The runtime writes the record as the run completes. Naming it is a small model call in a nameThread slot, made once on the first exchange. Both the list and the transcripts are keyed on the signed-in user, exactly as memory is.
Two limits. A message you send mid-answer is not in the replay: he takes it into account, but the record is written once per completed run and holds the question that started it. And naming can fail while the record does not: a provider having a bad minute, or a conversation he decides is not worth naming (a greeting, a test message), leaves it listed by its first question.
Update him
His definition ships inside the orchestrator binary, so upgrading the platform can bring a newer one. When it does, the page reports Update available. Press Roll out update to publish the new bundle as a version and roll his deployment onto it. A newer bundle waits for you, and installing never happens twice.
Put him back the way he shipped
There is one button on a deployed agent, and it always does the same thing: republish the definition this orchestrator ships and replace his pods with it. Only its name changes:
| It reads | When |
|---|---|
| Roll out update | A newer bundle ships with this orchestrator |
| Redeploy | The deployment is in a failed state |
| Reinstall from stock | Neither: he is current and running |
Reach for the last one after editing him into a state you would rather undo, or when he is running something you no longer trust.
This replaces changes you made to the agent, but it does not destroy them. The live definition is frozen as its own version first, tagged agent-edited-…, and only then overwritten. Watch for the Edited badge; the confirmation names what is about to happen.
An edit you want back is under Versions with that tag: read it, copy from it, or deploy it again. The tag is derived from the definition's content, so reinstalling twice from the same edited state reuses one version, and an unedited reinstall freezes nothing.
Trace him
He installs with tracing already on, so his runs are under Traces from the first one: every model call, tool call and block. Tracing is off by default for what you build because it costs throughput; a chat agent has none to lose, and he can deploy, so "what did he do, and was he told to" should be answerable about every run. Turn it off with the toggle if you would rather; the choice sticks across redeploys and roll-outs. Either way his pods are replaced, because the runtime reads the tracing setting when it starts.
Give him more turns
One answer is a loop: he calls a tool, reads what came back, and decides what to do next. Each pass is a turn, and there is a limit on how many one answer may take. Reaching it ends the run without an answer, and the panel says he ran out of steps.
He ships with a limit of 60, enough to read the traces for a failing deployment, work out why, write a fix and prove it with a test suite. Raise it under Admin → Platform agent if your tasks are longer, or lower it to cap what a stuck run can spend. Anything from 1 to 200 is accepted. Leave the field empty to use the limit his definition ships with; clearing it is how you get back to that after setting a number.
It is a ceiling and not a budget: an answer that takes four turns takes four. Saving it replaces his pods, the same as tracing and for the same reason.
Remove him
Press Remove to undeploy him. The integration is kept, so any changes you made survive and installing again reuses them. This is the reversible half.
To delete the integration as well, call the API with ?purge=1. That also removes his skill resources and the cluster secret holding the provider key. It is the irreversible half, which is why it is not a button.
His workspace
Dr. Octo can try a fix rather than describe one. He runs on the agentic runner, a heavier image than the one your integrations get, carrying a shell, curl, jq, the standalone octo CLI, dolphin, and a scratch directory at /workspace. That gives him a loop: write a definition and a _test.yaml suite beside it, run the suite, read the failures, fix, run again, and invoke a single flow to see what it returns.
The workspace is scratch: an emptyDir that dies with his pod, capped at 100Mi. Anything worth keeping he saves back through the API as a real integration, and he will tell you so. A flow he invokes runs standalone, with in-process queues, in-memory KV, no cluster resources and no connectors reaching anything he was not given; it proves the shape and logic of a definition, never what your deployment would see, and he is told to say which of the two he has shown you. Each command is killed after two minutes, so there is no octo run and nothing long-lived.
He is deployed on that runner automatically. On an installation whose chart does not configure it, Install is blocked with the reason shown, because his definition does not load without it.
Why the web search asks
web_search and web_extract ask you first because of what they bring back: a search result or a fetched page is text somebody else wrote and can carry instructions addressed to him. Everything else he can reach is either a GET or an orchestrator call whose effect is on this installation's own data, and stays free, because a panel that asked about every read would be a panel whose questions nobody reads. The objective and the queries (or the URLs) ride in the question, so you are allowing that call rather than the tool in general. If you do not answer, the call is denied when his clock runs out, and he answers from what he knows and says he could not check the web.
It is an ordinary ai-agent setting, authorize on the tool, so you can gate another tool of his the same way, or take this one off, by editing his definition and rolling out. Human in the Loop walks through building the same thing into an agent of your own.
Restrict what he can do
Set the limit on what he may do in his own definition rather than with a flag on the orchestrator. Open him in the editor, find the octo_api tool (a rest-dynamic block inside the platform_operator specialist, and the only unrestricted one in the file) and add either setting:
- type: rest-dynamic
settings:
connector: octo
method: body.method
path: body.path
# Empty by default. Either of these narrows what he may call.
allowMethods: [GET]
pathPrefix: /integrationsUse allowMethods: [GET] to make him read-only, or pathPrefix to confine him to one part of the API. Save and roll out. The connector's baseURL is a boundary he cannot cross whatever the block says: a dynamic path resolves against it and cannot redirect the call elsewhere.
His observability_api tool is a second, independent boundary on a different connector with a different baseURL, so narrowing one does nothing to the other. It ships with allowMethods: [GET], and that line matters: the observability API also serves PUT /settings/retention and POST /retention/run, so without it he could change the retention policy or run a sweep. The alerts_api tool on the same connector allows GET, POST, PUT and DELETE but is bounded by pathPrefix to /alerts, so he can create and edit watches and nothing else on that service.
What those tools spend is the token of whoever is chatting, not his pod's, so what he can read there is what that person can read in the Traces view — and an administrator asking him to change the retention policy would succeed where a monitor asking the same thing gets a refusal he reports as one. http_fetch reaches any address his pod can, but it carries no credential, so it reads nothing from a service that wants one.
The command tools are a different kind of thing
Everything above narrows an address. The workspace tools do not have one. Deleting a tool is the control that works. http_fetch, run_command, invoke_flow and the rest are blocks in a definition you can edit, exactly like octo_api. Remove the block, save, roll out, and the capability is gone.
The allow list on those blocks is a guardrail, not a boundary. It names the five programs he may run, but octo is on that list and octo invoke runs whatever definition it is pointed at, including one written into his workspace a moment earlier. So the boundary is the pod: its own environment, which for him means his provider key, and the runtime ServiceAccount, with no cluster credentials beyond leases. That blast radius is why the agentic runner is a per-deployment choice.
http_fetch is the new reach. Before it, his outbound access was two baseURLs and his model provider; now it is anything his pod can route to, including other integrations' internal Services and, on a cloud node, the instance metadata endpoint. He is restricted to http and https, which stops file:// reading his own pod, but nothing in his definition restricts hosts. If that matters on your cluster, a NetworkPolicy on his deployment is the control.
He reads integration definitions, resource contents, pod logs and, for whoever is asking, stored log messages and captured request bodies from every deployment. All of that is text other people wrote, so a malicious string in a flow definition, or in a log line an integration was made to emit, is an instruction he may see. What bounds the damage is that he acts as the person chatting: an instruction he picks up can do no more than they could. Narrow allowMethods to bound it further, delete the command tools if the loop above is not worth that reach, and leave tracing on so that "was he told to do that" is answerable after the fact.
What he is made of
Read this before you change him. The whole file is in the editor.
| Piece | What it does |
|---|---|
chat flow | An HTTP source on /chat with server-sent events enabled, so answers stream a token at a time |
ai-agent block | The coordinator: streaming, agentId: dr-octo, memory keyed to the signed-in user and the conversation, a 500k-token window and summarising compaction, and userMemory on. stopWhen reads the X-Agent-Stop header, which is how the panel's Stop reaches a run it is not connected to. Two more ai-agent blocks sit in its tool slots, see the two specialists |
nameThread slot | On the coordinator: on the opening exchange only, asks llm-fast for a name for the conversation. The engine writes it, see Naming a conversation. The record itself is the runtime's |
octo_read | Reads the orchestrator's API with allowMethods: [GET], so the coordinator can answer a question but not change anything |
integration_builder | The first specialist (see below): writes a definition and proves it |
platform_operator | The second specialist: everything that changes the installation, and everything about what already happened |
send_email_report | Sends an HTML report through the platform's own email settings; the sender identity is the installation's, never the model's |
web_search, web_extract | Search the open web, and read named pages, through a parallel connector, when the installation has a key. Without one they answer that search is not configured and say not to call again, which costs one turn rather than a retry loop. The two tools that ask you first |
navigate_to | Moves your browser to a page, when the answer to a question is a page |
check_expression | Evaluates one CEL expression, so he can prove one rather than assert it |
remember, forget, search_memory | Built into the runtime rather than written here. What he has already remembered about you is placed in every request, so there is no tool to call to recall it |
| Skills | Nine documents loaded on demand rather than sitting in every prompt. The coordinator holds seven (flow files, blocks, connectors, expressions, testing, the platform and email reports); the operator adds alerts, and the troubleshooter adds troubleshooting |
The two specialists
The coordinator keeps the cheap things (a read-only API call, a CEL check, a web search, navigation, the built-in memory tools, and the mail tool) and delegates the work to two ai-agent blocks standing in its tool slots, each with its own prompt, tools and loop. The coordinator writes a brief, the specialist works, and one report comes back as the tool result.
| Specialist | Holds | For |
|---|---|---|
integration_builder | read_workspace_file, write_workspace_file, list_workspace, invoke_flow, run_flow_tests, builder_check_expression, run_command, builder_octo_read | Writing and proving a definition: draft into the workspace, run the dolphin suite, read the report, fix, run again. Its API tool is GET-only, so nothing it produces is saved |
platform_operator | list_api_operations, read_api_docs, octo_api, list_observability_operations, read_observability_docs, observability_api, alerts_api, http_fetch, refresh_platform_ui | Changing and inspecting the installation: saving a definition, version tags, deploys, rollouts, scaling, secrets, and every question about stored logs and traces. With alerts_api and the alerts skill it also creates, previews and edits alert watches when you ask in chat. Its prompt carries a map of the API's routes, so it does not rediscover them from the OpenAPI on every call |
Each specialist keeps its own conversation while yours is going on: memoryThreadId: vars.toolScope with memoryVolatile: true, the scope the runtime mints for a tool branch. A second delegation in the same chat reaches the same builder with its earlier draft still in front of it, and none of it touches your conversation or the store that is backed up.
The split keeps each specialist's working context apart (YAML drafts and test reports in one, API responses in the other) and your conversation receives one report rather than forty tool results. It is also a boundary: the one unrestricted octo_api in the file is inside the operator, so a write to the platform is always a deliberate delegation. The cost is that a delegation is a whole model run, which is why the SSE route's maxDuration is 30 minutes and why the coordinator is told to answer from octo_read when that is the whole answer. The specialists cannot see your conversation, so a thin brief comes back as a specialist guessing confidently. Each specialist emits its tool_call and tool_result events to the same stream as the coordinator's, so the panel shows what it is doing.
He also triages alerts
A fourth flow, troubleshooter, is the same agent woken by an alert instead of by a question. Nobody is watching a panel, so it does not stream, and the two emails are the deliverable.
An alert reaches it because a watch's topic action publishes onto this deployment's own subject and the flow subscribes with an ordinary events source. The subject name has to match at both ends; ALERT_SUBJECT is that agreement, and alerts is the default. Whoever should hear about it rides on the action as reportTo. The flow's first block is a validate gate named only-bad-news: only notifications with kind open or repeat start a run, while resolve and close are logged and dropped, because the events source has no filter of its own.
It works in three acts. First it triages and reports: it reads the deployment's logs, compares the numbers against the baseline the alert carried, checks what was deployed and when, looks things up in these docs, and emails what it has before attempting anything. Second, if it can name the cause and the installation permits it, it tries a fix: scaling first, because it is reversible, and a definition change last, only through the builder so there is a test between the edit and the deploy. Third, it reports again with what was found, what was done and what is left.
It remembers what it has seen. The flow (not the agent) writes each episode into the persistent store under alert-history:<watchId>, capped at 50 episodes, and hands that history to the agent with every alert. That record separates a blip from a fault: once is transient and gets a report, four times in two hours is what the third act is for, and an entry saying a previous attempt did not hold stops it trying the same fix twice.
The third act is off by default, behind a checkbox on Admin → Agent: Allow Dr. Octo to troubleshoot applications. On means a model acting on your installation with nobody watching, started by an alert that fired at four in the morning. Off, he still triages, reads everything and reports; it only stops him carrying out what he recommends.
Saving it rolls the agent's pods, like the tracing toggle, because the runtime reads the setting at startup. Underneath it is the AGENT_TROUBLESHOOT_FIX binding, set the same way the model and the search key are.
The second report is sent by the flow rather than by the agent, so a run that used up its turns or gave up cannot end in silence after the alert has been marked as handled. It carries an appendix of what happened, assembled from the agent's own event stream: one line per turn, tool call and result, without the token-level events. The same lines are logged as they happen, so a run in progress is visible in the Logs view.
Two things to know before turning the third act on. Enabling tracing is a rollout: the pods restart, discarding the in-flight state, and it backfills nothing, so the skill tells him to reach for the logs first and to say in the report when he did this. And the alert body is evidence, not instruction: a watch name is text a person wrote, arriving in a prompt belonging to an agent that can reach the operator. The skill says so plainly, and the same goes for a log line or a page of these docs.
The self-healing loop maps the parts he sits between, and Agentic self-healing walks through creating the watch, pointing it at him, and putting a person in the loop over Slack.
What he remembers
He uses the runtime's agent memory with agentId: dr-octo and userMemory on, so none of this is written in his definition. Three things are stored:
| What it is | |
|---|---|
| Working memory | The transcript he re-reads. Compacted as it grows: memoryCompaction: summarize folds the early turns into a summary of them |
| Conversation history | What you can come back to. Never compacted, so it is what was actually said rather than the shortened version he still carries |
| Facts | What outlives a conversation: how you like flows written, which integration is yours, what to call you |
All three live in the orchestrator's own tables, keyed on the integration, so they survive a redeploy and a reinstall.
Facts are yours: they are keyed on the user the platform authenticated, which he can neither see nor set, because userId is resolved by the flow from an identity the chat route writes server-side. What he has remembered is placed in every request; past a bound the rest stay reachable through search_memory. Remembering under a name that already exists overwrites it. Tell him to forget something and he will, but a fact he was told not to keep may still be in the transcript; erasing the conversation from the memory viewer removes both, and so does purging the agent.
Run him yourself
He is a normal octo app, so run him standalone:
# llm-anthropic | llm-openai | llm-gemini | llm-openrouter
export LLM_CONNECTOR_TYPE=llm-anthropic
export LLM_API_KEY=sk-ant-...
export LLM_MODEL=claude-sonnet-4-6
export ORCHESTRATOR_URL=http://localhost:8090
# Optional. Without it his observability tools answer that the access was not granted.
export OBSERVABILITY_URL=http://localhost:8091
# Required. His workspace and the five programs he drives default to the agentic
# runner's paths, which do not exist here. A cli-run allow list is resolved when
# the flow is BUILT, so an entry pointing at a missing file fails the whole config
# on load rather than the first call that would have used it.
export AGENT_WORKSPACE=$(mktemp -d)
export AGENT_OCTO_BIN=$PWD/bin/octo
export AGENT_DOLPHIN_BIN=$PWD/bin/dolphin
export AGENT_CURL_BIN=$(command -v curl)
export AGENT_FIND_BIN=$(command -v find)
export AGENT_JQ_BIN=$(command -v jq)
bin/octo run --config orchestrator/agentThat build-time resolution is why his definition names every program through a variable rather than writing the path in the allow list. Change one to a literal and he stops loading anywhere but the cluster, including in the dolphin suite that ships beside him.
curl -N -X POST localhost:8080/chat \
-d '{"threadId":"t1","message":"list my integrations"}'Work this way when you are iterating on his prompt or a tool; it is far faster than rolling out to see the result.