Octov0.11.7
Testing

Running Tests

Run a suite from a terminal, the whole battery from the editor, and the same files in CI, with exit codes a pipeline can act on.

A suite is a file, so everything runs it: dolphin test in a terminal, the editor's Testing tab, an agent over MCP, and CI. They all shell out to the same binary and give the same verdict.

From a terminal

dolphin test ./flows              # every *_test.yaml in the directory, against it
dolphin test flows/orders.yaml    # its companion orders_test.yaml, against that file
dolphin test flows/orders_test.yaml   # the suite, against the flow file beside it

test is the default subcommand, so dolphin ./flows works too. Paths may come before or after the flags, the way go test accepts them.

A suite is paired with the config the way Go pairs a test file with its package. Point at a directory and every *_test.yaml directly inside it runs against that directory (subdirectories are not walked). Point at a flow file and its companion runs against that one file. Use --config when the suite does not live beside the flows it exercises:

dolphin test --config flows/orders.yaml tests/smoke_test.yaml

Every named suite is loaded and validated before any case runs, so a file dolphin refuses fails the run immediately.

Flags

FlagDefaultMeaning
--config <path>the flows beside the suiteThe flows to test against.
--env-file <path>noneA .env file every case runs with. A missing file is a hard error.
--traces-dir <path>noneRun every case under tracing and write one .trace.jsonl per case.
--junit <path>noneWrite a JUnit XML report.
--report-json <path>noneWrite a machine-readable JSON report.
--parallel <n>one per CPUHow many cases to run at once.
--fail-fastoffStop after the first failing case.
-voffName every case, and let octo's logs through.

Flags accept one or two dashes.

Tracing a suite

--traces-dir runs each case with tracing on and writes its trace beside the others:

dolphin test ./flows --traces-dir ./traces

Each file records the message entering and leaving every block, by address. A suite is the ideal thing to trace: the cases already say how to exercise the flow, and the mocks they carry mean nothing real is called — so it is a way to see the true shape of a flow's messages without touching production.

One file per case rather than one per run, because cases run in parallel and two processes appending to one file would interleave into neither.

Bodies and variables are captured by default and can carry credentials and personal data. Set OCTO_TRACING_BODIES=false and OCTO_TRACING_VARS=false to record the sequence only.

Finding octo

dolphin runs your tests by running the real octo. It stops at the first of:

  1. $OCTO_PATH, the binary itself or a directory holding it
  2. ./octo, in the current directory
  3. octo, on your PATH

$OCTO_PATH is an override, not a hint: when it is set and does not name a runnable octo, dolphin fails rather than testing against a different build.

The environment

A suite's env: wins over --env-file, which wins over what you have exported, so a test cannot quietly use the real API key in your shell. For a value several suites share, put it in a file:

dolphin test ./flows --env-file .env.test

A test key is fake by construction, so it can be committed. This repository keeps one at samples/.env.test.

From the editor

The Testing tab's toolbar runs the open suite; the header's Run tests runs every suite in the document at once, grouping the results one block per suite. It stages the config the suite is for and runs the same dolphin binary.

It holds back suites it knows dolphin would refuse, and names them; see running every suite. A suite held back there still runs in a terminal and in CI.

The dev .env is deliberately not injected into a test run, unlike a debug run. It would make the tab and dolphin test disagree, and would let a "test" authenticate with your real credentials.

In CI

.github/workflows/validate.yml
- run: task runtime:build
- run: task runtime:test:samples

Or directly, for a project that keeps its flows in one directory:

.github/workflows/test.yml
name: test
on: [push, pull_request]

jobs:
  flows:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-go@v5
        with: { go-version: stable }

      - name: Install octo and dolphin
        run: |
          go install github.com/juancavallotti/octo/runtime/octo@latest
          go install github.com/juancavallotti/octo/runtime/dolphin@latest

      - name: Test the flows
        run: dolphin test ./flows --env-file .env.test --junit report.xml

      - name: Publish the report
        if: always()          # a failing run is exactly when you want the report
        uses: actions/upload-artifact@v4
        with:
          name: flow-tests
          path: report.xml

--junit writes one <testsuite> per file and one <testcase> per case. A failed case and an errored one are tagged differently (assertion versus run), so the exit-code distinction below survives into the dashboard.

This repository tests all of its own samples this way.

Exit codes

CodeMeaning
0Every case passed.
1A case ran and did not do what it said. The flows are wrong.
2A case never ran: a block address does not resolve, the config does not parse, there is no usable octo. The suite is wrong.

An errored case (one that never ran) produces exit 2, not 1, even when other cases in the same file failed normally.

A machine-readable report

--junit is for a CI system's test UI. --report-json <path> is for a program:

dolphin test ./flows --report-json report.json

It carries the tally, each case's failures with their detail, and, for every case that ran, the outcome itself:

{
  "dolphin": "dolphin 0.6.2",
  "wallMs": 412,
  "workDir": "/tmp/dolphin-123",  // only when something failed and was kept
  "totals": { "cases": 5, "passed": 4, "failed": 1, "errored": 0,
              "skipped": 0, "notRun": 0, "elapsedMs": 380 },
  "suites": [{
    "path": "flows/orders_test.yaml", "config": "flows/orders.yaml",
    "flow": "orders", "elapsedMs": 380,
    "cases": [{
      "name": "a large order is declined",
      "status": "passed | failed | errored | skipped | not-run",
      "elapsedMs": 80,
      "summary": "one line",
      "failures": [{ "summary": "…", "detail": "…multi-line…" }],
      "reproduce": "octo invoke …",
      // What the flow actually did, the other half of a failure. Present on
      // every case that RAN, not only the ones that failed.
      "outcome": { "result": { }, "dropped": false, "error": "…", "spies": { } }
    }]
  }]
}

Empty fields are omitted, so a passing case carries no failures and a skipped case carries no outcome. Three things to rely on:

  • A report is written for exit 0, 1 and 2. Only a run that never got as far as running (no usable octo, a config that will not parse) has nothing to write.
  • status is a string, never an index, so an absent field and "not run" cannot both read as 0.
  • totals.elapsedMs sums the cases; wallMs is the wall clock, which is shorter under --parallel.

There is deliberately no schema version while octo is pre-release; the document grows additively.

The editor's Testing tab reads this report rather than reimplementing the assertions.

Parallelism and isolation

One case is one octo process. Mocks and spies are baked into the flow tree when it is built, so a service can only serve the case it was built for. Cases are fully isolated, and dolphin runs them in parallel, one per CPU. --parallel 1 serializes them, which you want when the config holds a connector that binds a port.

Mocking a block does not stop its connector from starting. A suite over a config with a database connector still opens the database, so its DSN has to be valid even when every SQL block is mocked. file::memory: is usually the answer.

See also

On this page