The Platform MCP Server
Build and operate integrations with coding agents through the platform MCP server.
The platform is itself an MCP server: it exposes the whole authoring lifecycle (read the runtime schema, draft and validate YAML, save integrations, run flows, read logs) as MCP tools. Connect a coding agent like Claude Code and it can build, test, and debug integrations against your live platform.
How to connect
The platform serves MCP at /mcp on its public origin (streamable HTTP;
SSE is not supported). The endpoint is an OAuth 2.1 protected resource:
claude mcp add --transport http octo https://<your-platform-host>/mcpThe first request without a token gets a 401 whose WWW-Authenticate header
points at /.well-known/oauth-protected-resource/mcp (RFC 9728); the client
discovers the authorization server there, signs you in, and retries with a
bearer access token, the handshake described in
Secure an MCP Server with OAuth. That handshake is the only way
in: /mcp verifies access tokens the identity provider minted, and nothing else.
One-shot calls (invoke_flow, evaluate_cel, run_tests) are isolated per
MCP session, so concurrent clients never step on each other. A long-running
dev run belongs to you and one integration, so a second
session running the same integration attaches to the run already up.
The standalone app exposes the same MCP handler, unauthenticated, at
http://localhost:<port>/mcp. Every tool below works the same way there,
backed by the local disk store.
What the server exposes
Authoring tools
| Tool | What it does |
|---|---|
list_integrations | Every saved integration as { id, name }. |
open_integration | One integration by id: { id, name, definition }, the runtime YAML. |
list_flows | The flows in an integration as { name, source } (source is the trigger type, or null for a sourceless flow). |
validate_definition | Validate draft YAML against the runtime schema without saving; returns descriptive errors. |
create_integration | Create an integration from a name and a YAML definition. |
update_integration | Overwrite an integration's whole definition (and optionally rename it). |
Flow tools
An integration is one YAML file holding many flows. These edit one of them
and leave every other byte of the file as it was, where update_integration
would have you reproduce the whole file from memory. Each mutating tool
validates the spliced result before saving, so an edit that would leave the
integration unloadable is refused.
| Tool | What it does |
|---|---|
get_flow | One flow's YAML, exactly as it appears in the file. Hand it back to update_flow to edit it. |
add_flow | Append a new flow. Errors when the name is already taken. |
update_flow | Replace a flow in place. Giving the replacement a different name renames it where it sits. |
delete_flow | Remove a flow by name. |
In each, definition is the YAML of that flow alone, a mapping starting
with name:, not the whole file and not a flows: list.
Run and debug tools
| Tool | What it does |
|---|---|
can_start_integration | Pre-flight: is a runner available, does the definition validate? |
run_integration | Start (or restart) the integration as a dev run, a pod of its own; returns its stable public test URL for networked integrations. Runs the saved definition, and refuses ad-hoc env (put values in the integration's .env.dev resource). |
invoke_flow | Run one named flow once, from a saved id or an inline definition, with optional JSON data, vars and env, without starting sources. Returns { ok, timedOut, dropped, output, logs }, where output is the flow's result message ({event_id, variables, body}) as JSON text. Takes breakAt, spies and mocks; see below. |
evaluate_cel | Evaluate a single CEL expression against sample body/vars/env without running a flow. |
get_run_logs | The run's output for an integration id, as plain text. |
stop_integration | Stop the run for an integration id. |
The three run-control tools take the integration id because a run belongs to
(you, that integration) rather than to your MCP session. A run you start is
the same run the editor attaches to, so an agent can debug what a human is
looking at. It runs what is saved, so an agent must save an edited flow
(update_flow does) before the running app reflects it. The one-shot tools
(invoke_flow, evaluate_cel, run_tests) run the definition in the
request. See Dev runs.
Debugging a flow with invoke_flow
Three features compose on one run, each addressing a block by the same path
grammar: <flow>.<block>, descending into a composite with a bracketed branch
(orders.charge, orders.checkHeader[else].api-call, orders[error].notify).
| Feature | Behaviour |
|---|---|
breakAt | Run until this block, then stop. The result carries breakpoint: { reached, block, message, error }. reached: false means the flow took another branch: a normal outcome, not an error. |
spies | Record every message crossing these blocks without changing what the flow does. The result carries spies: [{ address, records }], where each record shows what the block received and what it produced, including the two outcomes that are not a message: it dropped, or it failed. seq orders records across all spies, so a multi-spy trace reads as one timeline. |
mocks | Stand in for these blocks so the real one never runs. This is how a flow whose blocks call a payment API or an LLM gets exercised without one. |
Two rules the runtime enforces. A mock replaces its target, so an unmatched
message fails the block rather than falling through to the real one; give a
default, or a trailing case with when: "true", unless you mean an unmatched
message to be an error. Each case does exactly one thing (body, error,
or drop); vars only goes alongside a body.
{
"id": "orders-api",
"flow": "checkout",
"data": "{\"amount\": 250}",
"spies": ["checkout.validate"],
"mocks": {
"checkout.charge": {
"cases": [{ "when": "body.amount > 100", "error": "card declined" }],
"default": { "body": { "charged": true } }
}
}
}You cannot breakAt or spy a block inside a mocked block: the mock deleted
that subtree, and the request is rejected rather than silently reporting
nothing.
Documentation tools
| Tool | What it does |
|---|---|
getSchema | No argument: a compact index of every block and connector type. With elementName (e.g. "log", "http"): that element's full spec, including its fields and, for a composite, the addressBranches a breakpoint/spy/mock address needs to descend into it (handle-errors[process].<block>). |
getExamples | No argument: an index of worked example integrations. With a slug: that example's full runtime YAML. |
getCelFunctions | The CEL variables and functions in scope for message expressions, Octo's own plus the extension libraries; pass a functionName for one entry, or a library (strings, lists, encoders, math, comprehensions, sets, regex, optional) for one family. |
These exist as tools (not only resources) because some MCP clients surface only tools.
Resource tools
| Tool | What it does |
|---|---|
list_env_keys | Every env var an integration expects, as { name, default, required, source }: declared vars plus keys found in env resource files. |
list_resources | An integration's stored resources as { id, kind, name } (kind is env or template). |
open_resource | Read one resource's content. |
create_resource | Create an env file or a template on an integration. |
update_resource | Update a resource's content, kind, or name. |
delete_resource | Delete a resource. |
Editor tools
The editor's own bookkeeping: the test inputs, mocks and spies the canvas shows. An agent that has just debugged a flow can leave the setup behind, so the user opens the canvas and finds the mocks placed and the ▶ menu ready to run.
These are not tests. Nothing here is an assertion or is read by
dolphin test, by CI, or by a deployed runtime; that is what the
test suite tools below are for. An agent asked to "add tests" that
reaches for set_mock has configured the editor rather than written a test.
| Tool | What it does |
|---|---|
list_block_addresses | Every block address in an integration, grouped by flow. Call this before anything that takes an address. |
get_flow_meta | One flow's saved { inputs, mocks, spies }. |
set_test_input / delete_test_input | A named message the flow can be run with from the ▶ menu. |
set_mock / delete_mock | Stand in for a block on every run the editor makes. |
set_spy / delete_spy | Record what crosses a block. |
Compose no address by hand. A block only has one when its name is unique
among its siblings and free of ., [ and ]; list_block_addresses lists
exactly the ones that exist. An address naming nothing is refused, because a
mock keyed by one looks placed and silently never fires; give the block a
distinct name with update_flow, which these tools will not do for you.
Renaming a flow with update_flow moves its bookkeeping with it, addresses
included.
Tests
A dolphin test suite is committed, CI runs it, and
dolphin test in a terminal gives the same verdict.
| Tool | What it does |
|---|---|
list_test_suites | The suites stored for an integration, with their case names. |
get_test_suite | One flow's suite as YAML, exactly as stored, comments and all. |
set_test_suite | Write a whole suite byte for byte. The flow under test is the file's own flow: key. |
set_test_case | Add or replace one case, structurally. Rewrites the file, so comments are lost; use set_test_suite when they matter. |
delete_test_case | Remove one case by name. |
run_tests | Run the suites and report what happened. |
Testing with an agent walks a whole session through these.
run_tests's ok means a report came back, not that the tests passed. The
verdict is in totals. A failed case is a flow that did not do what the
case says; an errored case is one dolphin could not run, so the test is
wrong. A case that did not pass also reports the message the flow produced.
Anything that would stop dolphin loading the file is refused before it is
written; dolphin rejects unknown keys, so a misspelled spys: takes the
whole file down rather than going quietly.
Values in a suite, JSON text in the meta file. A test suite holds bodies and variables as values, and so do the suite tools. The editor's bookkeeping stores them as JSON text, but its tools still take and return values and convert for you. Neither side ever wants a string of JSON.
MCP resources and prompts
The same catalogues are also published as resources: octo://runtime/schema
(the capability catalogue, generated from the runner binary, so it reflects
exactly what your deployment supports), octo://examples (the example index),
and octo://examples/<slug> (each example's YAML). Two prompts,
create-integration and write-effective-integrations, give a connected agent
step-by-step authoring guidance and design best practices.
The authoring loop
Ask Claude Code to "build me an integration that posts a daily summary to Slack" and it runs this loop:
Orient. The agent calls list_integrations to see what exists, and
open_integration to read anything related.
Learn the vocabulary. It calls getSchema for the index of block and
connector types, getSchema("slack-send-message") (and friends) for exact
fields, and getExamples to adapt a worked example instead of guessing YAML
syntax.
Draft and validate. It writes a definition and calls validate_definition
until the errors are gone: no save, no side effects.
Save. create_integration (or update_integration on iteration) persists
the definition. list_env_keys reveals which env vars still need values; the
agent asks you for secrets and stores them with create_resource as an env
file, which is where a dev run reads its environment.
Test one flow fast. invoke_flow runs a single flow with sample data, no
sources started, and returns the output and logs in one call: tweak, invoke,
read, repeat. evaluate_cel checks a gnarly expression in isolation.
Run for real. run_integration boots the integration as a dev run and
returns its test URL when it serves HTTP. The agent exercises the endpoints,
reads get_run_logs for the runtime's load errors and log output, fixes,
saves (the run reads the stored definition), and re-runs.
stop_integration tears it down.
Validation is best-effort: a clean validate_definition does not guarantee the
runtime loads the definition, and the pre-flight can flag YAML the runtime
accepts. The runtime is the final judge, which is why the loop ends at
get_run_logs.
Related pages
- Expose an MCP Server: make your own integrations MCP servers with
mcp-router. - Secure an MCP Server with OAuth: the same RFC 9728 handshake the platform endpoint uses, applied to your integrations.
- Capstone: A Production Slack Agent: the kind of integration a connected coding agent can build and debug through these tools.
- Testing a flow: the suite format
set_test_caseandrun_testswrite and run. - Testing flows in the editor: the same suites, in the Testing tab a user opens afterwards.