Testing with an Agent
Have an agent write and run a flow's dolphin suite over MCP, and know which tools touch the committed file.
Writing test cases suits an agent: the work is repetitive, the block addresses have to
be exactly right, and the feedback loop is a runner that says pass or fail. octo's
MCP server exposes both halves, writing a suite and running it.
Everything here writes the same <flow>_test.yaml a human writes; there is no
agent-specific test format.
First, the distinction that costs the most
Two families of tools have similar names and write to different places:
| Tool family | Writes | Read by |
|---|---|---|
set_test_suite, set_test_case, delete_test_case, list_test_suites, get_test_suite, run_tests | the committed <flow>_test.yaml | dolphin test, CI, the Testing tab, a human reviewer |
set_test_input, set_mock, set_spy (and their deletes) | the editor's scratch bookkeeping | the canvas and its ▶ menu, not a deployed runtime and not dolphin test |
The bookkeeping tools let an agent that has just debugged a flow leave the mocks placed on the canvas. Nothing in that file is an assertion, and nothing in it survives into CI. The suite tools leave behind something that outlives the session.
An agent asked to "add tests" that reaches for set_mock has configured the editor,
not written a test. The mock will be on the canvas and absent from every pipeline.
A session, end to end
Find the flow and its addresses
open_integration { id: "…" }
list_block_addresses { id: "…", flow: "checkout" }Never compose an address by hand. A block only has one when its name is unique
among its siblings and free of ., [ and ]. list_block_addresses returns exactly
the addresses that exist, and a mock keyed by one that does not is refused rather than
written.
Write the suite
set_test_suite { id: "…", content: "<the whole YAML>" }The flow under test is the file's own flow: key. The content is stored byte for
byte; this is the one write that preserves comments. Anything that would stop
dolphin loading the file is refused before anything is written, including unknown
keys: a misspelled spys: takes the whole file down.
To amend an existing suite one case at a time:
get_test_suite { id: "…", flow: "checkout" } → the YAML, as stored
set_test_case { id: "…", flow: "checkout", case: { … } }
delete_test_case { id: "…", flow: "checkout", name: "…" }set_test_case rewrites the file structurally, so its comments are lost. It says so
when it does. To keep them, read the file with get_test_suite, edit the text, and
write it back with set_test_suite.
Bodies and variables are values, not JSON strings. data: { amount: 10 }, never
data: "{\"amount\": 10}". The editor's bookkeeping stores them as JSON text, but even
those tools take and return values and convert for you.
Run them
run_tests { id: "…" } → every suite
run_tests { id: "…", flow: "checkout" } → oneok means a report came back, not that the tests passed. dolphin exits non-zero
for a failing case. The verdict is in totals.
Read totals, then act on which counter moved:
| Counter | What it means | What to fix |
|---|---|---|
failed | The flow ran and did not do what the case says | The flow, or the expectation |
errored | The case never ran: a bad address, an undeclared input | The test |
skipped | A skip: reason on the case | Nothing; it was deliberate |
A case that did not pass also carries outcome, what the flow actually produced,
which is usually enough to fix the case without another round trip.
Things an agent gets wrong
Asserting an exact body on a result carrying a generated id or a timestamp passes
once and fails forever after. Use that: with a CEL expression over the stable part.
Mocking a composite and then asserting on what ran inside it cannot work: the mock replaces the block and its whole subtree. Mock a leaf within it.
Writing a mock with no default means a message matching no case fails the block; it
does not fall through to the real one, which is gone.
Forgetting env: breaks the build of the flow: a config's environment resolves when it
loads, so a flow whose connector reads ${SOME_KEY} cannot be built without one, and
mocking the block that uses it does not help.
What the run environment is
run_tests stages the integration's config, its suites and its declared resources into
a scratch directory and runs the real dolphin binary against them, one case at a
time.
Two deliberate differences from an interactive run. The dev .env is not injected,
because that would make the tool and dolphin test disagree and would let a "test"
authenticate with real credentials; a suite states what it needs in its own env:.
Cases run serially, because each case starts a full service with its non-source
connectors, and parallel runs would contend for the same ports.
The reproduce command dolphin prints for a failing case is stripped from the response: it names a directory that has been deleted, and would leak the server's filesystem layout.
Afterwards
The suite is a normal file in the integration. Whoever opens the editor next sees it in the Testing tab, can run it there, and gets the same verdict:

See also
- MCP server: the full tool catalogue, including the ones that build the flow in the first place.
- Test file reference: the schema a
set_test_suitepayload has to satisfy. - Writing test cases: the same format, explained for people.