Running Tests
Run a suite from a terminal, the whole battery from the editor, and the same files in CI, with exit codes a pipeline can act on.
A suite is a file, so everything runs it: dolphin test in a terminal, the editor's
Testing tab, an agent over MCP, and CI. They all shell out to the same binary and give
the same verdict.
From a terminal
dolphin test ./flows # every *_test.yaml in the directory, against it
dolphin test flows/orders.yaml # its companion orders_test.yaml, against that file
dolphin test flows/orders_test.yaml # the suite, against the flow file beside ittest is the default subcommand, so dolphin ./flows works too. Paths may come
before or after the flags, the way go test accepts them.
A suite is paired with the config the way Go pairs a test file with its package.
Point at a directory and every *_test.yaml directly inside it runs against that
directory (subdirectories are not walked). Point at a flow file and its companion runs
against that one file. Use --config when the suite does not live beside the flows it
exercises:
dolphin test --config flows/orders.yaml tests/smoke_test.yamlEvery named suite is loaded and validated before any case runs, so a file dolphin refuses fails the run immediately.
Flags
| Flag | Default | Meaning |
|---|---|---|
--config <path> | the flows beside the suite | The flows to test against. |
--env-file <path> | none | A .env file every case runs with. A missing file is a hard error. |
--traces-dir <path> | none | Run every case under tracing and write one .trace.jsonl per case. |
--junit <path> | none | Write a JUnit XML report. |
--report-json <path> | none | Write a machine-readable JSON report. |
--parallel <n> | one per CPU | How many cases to run at once. |
--fail-fast | off | Stop after the first failing case. |
-v | off | Name every case, and let octo's logs through. |
Flags accept one or two dashes.
Tracing a suite
--traces-dir runs each case with tracing on and writes its
trace beside the others:
dolphin test ./flows --traces-dir ./tracesEach file records the message entering and leaving every block, by address. A suite is the ideal thing to trace: the cases already say how to exercise the flow, and the mocks they carry mean nothing real is called — so it is a way to see the true shape of a flow's messages without touching production.
One file per case rather than one per run, because cases run in parallel and two processes appending to one file would interleave into neither.
Bodies and variables are captured by default and can carry credentials and personal
data. Set OCTO_TRACING_BODIES=false and OCTO_TRACING_VARS=false to record the
sequence only.
Finding octo
dolphin runs your tests by running the real octo. It stops at the first of:
$OCTO_PATH, the binary itself or a directory holding it./octo, in the current directoryocto, on yourPATH
$OCTO_PATH is an override, not a hint: when it is set and does not name a
runnable octo, dolphin fails rather than testing against a different build.
The environment
A suite's env: wins over --env-file, which wins over what you have exported, so a
test cannot quietly use the real API key in your shell. For a value several suites
share, put it in a file:
dolphin test ./flows --env-file .env.testA test key is fake by construction, so it can be committed. This repository keeps one
at samples/.env.test.
From the editor
The Testing tab's toolbar runs the open suite; the header's Run tests runs every
suite in the document at once, grouping the results one block per suite. It stages the
config the suite is for and runs the same dolphin binary.
It holds back suites it knows dolphin would refuse, and names them; see running every suite. A suite held back there still runs in a terminal and in CI.
The dev .env is deliberately not injected into a test run, unlike a debug run.
It would make the tab and dolphin test disagree, and would let a "test" authenticate
with your real credentials.
In CI
- run: task runtime:build
- run: task runtime:test:samplesOr directly, for a project that keeps its flows in one directory:
name: test
on: [push, pull_request]
jobs:
flows:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with: { go-version: stable }
- name: Install octo and dolphin
run: |
go install github.com/juancavallotti/octo/runtime/octo@latest
go install github.com/juancavallotti/octo/runtime/dolphin@latest
- name: Test the flows
run: dolphin test ./flows --env-file .env.test --junit report.xml
- name: Publish the report
if: always() # a failing run is exactly when you want the report
uses: actions/upload-artifact@v4
with:
name: flow-tests
path: report.xml--junit writes one <testsuite> per file and one <testcase> per case. A failed
case and an errored one are tagged differently (assertion versus run), so the
exit-code distinction below survives into the dashboard.
This repository tests all of its own samples this way.
Exit codes
| Code | Meaning |
|---|---|
| 0 | Every case passed. |
| 1 | A case ran and did not do what it said. The flows are wrong. |
| 2 | A case never ran: a block address does not resolve, the config does not parse, there is no usable octo. The suite is wrong. |
An errored case (one that never ran) produces exit 2, not 1, even when other cases in the same file failed normally.
A machine-readable report
--junit is for a CI system's test UI. --report-json <path> is for a program:
dolphin test ./flows --report-json report.jsonIt carries the tally, each case's failures with their detail, and, for every case that ran, the outcome itself:
{
"dolphin": "dolphin 0.6.2",
"wallMs": 412,
"workDir": "/tmp/dolphin-123", // only when something failed and was kept
"totals": { "cases": 5, "passed": 4, "failed": 1, "errored": 0,
"skipped": 0, "notRun": 0, "elapsedMs": 380 },
"suites": [{
"path": "flows/orders_test.yaml", "config": "flows/orders.yaml",
"flow": "orders", "elapsedMs": 380,
"cases": [{
"name": "a large order is declined",
"status": "passed | failed | errored | skipped | not-run",
"elapsedMs": 80,
"summary": "one line",
"failures": [{ "summary": "…", "detail": "…multi-line…" }],
"reproduce": "octo invoke …",
// What the flow actually did, the other half of a failure. Present on
// every case that RAN, not only the ones that failed.
"outcome": { "result": { }, "dropped": false, "error": "…", "spies": { } }
}]
}]
}Empty fields are omitted, so a passing case carries no failures and a skipped case
carries no outcome. Three things to rely on:
- A report is written for exit 0, 1 and 2. Only a run that never got as far as
running (no usable
octo, a config that will not parse) has nothing to write. statusis a string, never an index, so an absent field and "not run" cannot both read as0.totals.elapsedMssums the cases;wallMsis the wall clock, which is shorter under--parallel.
There is deliberately no schema version while octo is pre-release; the document grows additively.
The editor's Testing tab reads this report rather than reimplementing the assertions.
Parallelism and isolation
One case is one octo process. Mocks and spies are baked into the flow tree when
it is built, so a service can only serve the case it was built for. Cases are fully
isolated, and dolphin runs them in parallel, one per CPU. --parallel 1 serializes
them, which you want when the config holds a connector that binds a port.
Mocking a block does not stop its connector from starting. A suite over a config
with a database connector still opens the database, so its DSN has to be valid even
when every SQL block is mocked. file::memory: is usually the answer.
See also
- Writing test cases: the suite format, through worked examples.
- Test file reference: every key, field by field.
- Testing in the editor: the same suites, in a form.