LLM Providers
The OpenAI, Anthropic, Gemini, and OpenRouter LLM connectors.
Octo ships four LLM provider connectors: llm-openai, llm-anthropic, llm-gemini, and llm-openrouter. All four register under the connector category llm and expose the same provider-agnostic completion capability, so the AI blocks (ai-agent, ai-router, ai-mapping, ai-retry) reference any of them interchangeably through their connector setting.
Three of them talk to one vendor each. llm-openrouter talks to OpenRouter, which fronts hundreds of models from many vendors behind one key, so reach for it to try several models without an account per vendor.
All four are service connectors and provide no source. The blocks that bind to them are the AI blocks, which accept any connector in the llm category; see AI in Octo for usage.
Each connector holds the provider policy (API key, model, and the default response token cap) and validates it at startup, so a missing key fails fast rather than on first request. An AI block's own maxTokens overrides the connector default per request.
llm-openai connector
A configured OpenAI Responses API client.
| Setting | Type | Required | Default | Description |
|---|---|---|---|---|
apiKey | string | Yes | none | Authenticates with the OpenAI API; source from ${OPENAI_API_KEY}. Never logged. |
model | string | No | gpt-5.4 | Model ID. |
maxTokens | int | No | model default | Default response token cap (0 = the model default); a request may override it. |
reasoning | enum default | none | minimal | low | medium | high | No | default | How much reasoning effort a reasoning-capable model spends. default sends no reasoning field at all and lets the model choose; the rest are sent as written. Any effort other than none also asks for a reasoning summary, which is what makes thinking visible. |
baseURL | string | No | OpenAI API | Overrides the API endpoint, for proxies, Azure, or OpenAI-compatible servers (any endpoint speaking the Responses protocol). |
llm-anthropic connector
A configured Anthropic Messages API client.
| Setting | Type | Required | Default | Description |
|---|---|---|---|---|
apiKey | string | Yes | none | Authenticates with the Anthropic API; source from ${ANTHROPIC_API_KEY}. Never logged. |
model | string | No | claude-sonnet-4-6 | Model ID. |
maxTokens | int | No | 4096 | Default response token cap; a request may override it. |
thinking | enum off | adaptive | budgeted | No | off | Extended thinking. adaptive lets the model decide how much; budgeted caps it at thinkingBudget. |
thinkingBudget | int | No | none | Thinking token budget, used when thinking: budgeted. Must be at least 1024 and below maxTokens, which it counts towards. Validated at startup and re-checked per request: a block whose own maxTokens is too small to hold the budget omits thinking for that call rather than failing it. |
baseURL | string | No | Anthropic API | Overrides the API endpoint (for proxies or testing). |
llm-gemini connector
A configured Google Gemini client.
| Setting | Type | Required | Default | Description |
|---|---|---|---|---|
apiKey | string | Yes | none | Authenticates with the Gemini API; source from ${GEMINI_API_KEY}. Never logged. |
model | string | No | gemini-3.5-flash | Model ID. |
maxTokens | int | No | model default | Default response token cap (0 = the model default); a request may override it. |
thinking | enum off | dynamic | budgeted | No | off | Thinking. dynamic lets the model decide how much; budgeted caps it at thinkingBudget. Either also asks for thought content to be returned. |
thinkingBudget | int | No | none | Thinking token budget, used when thinking: budgeted. |
baseURL | string | No | Gemini API | Overrides the API endpoint (for proxies or testing). |
llm-openrouter connector
A configured OpenRouter client, speaking Chat Completions.
Model ids here are vendor-prefixed as OpenRouter publishes them (anthropic/claude-sonnet-4.5, openai/gpt-5.4, google/gemini-3.5-flash) and a variant suffix (:free, :thinking) selects a routing variant. There is no bare model name.
| Setting | Type | Required | Default | Description |
|---|---|---|---|---|
apiKey | string | Yes | none | Authenticates with the OpenRouter API; source from ${OPENROUTER_API_KEY}. Never logged. |
model | string | No | anthropic/claude-sonnet-4.5 | Vendor-prefixed model id. |
maxTokens | int | No | model default | Default response token cap (0 = the model default); a request may override it. |
reasoning | enum default | none | minimal | low | medium | high | No | default | How much reasoning effort a reasoning-capable model spends. default sends no reasoning field at all and lets the model choose; none switches reasoning off explicitly. |
baseURL | string | No | OpenRouter API | Overrides the API endpoint (for proxies or testing). |
appName | string | No | none | Sent as X-Title. How OpenRouter attributes this traffic in its dashboards. |
siteURL | string | No | none | Sent as HTTP-Referer, for the same purpose. |
Traced calls through OpenRouter report what they actually cost. Every request asks for usage accounting, and OpenRouter answers with the amount it billed, including the per-request and per-image charges a token count cannot reconstruct. See Traces for how that reaches a trace.
Traced OpenRouter calls report the provider OPENROUTER, not the vendor behind whichever model answered. The vendor family names the token accounting, and OpenRouter reports an input count that already includes the tokens served from cache (the OpenAI-compatible convention), even when an Anthropic model behind it would have reported the uncached remainder on its own API. Pricing keys on that, so getting it wrong would mis-bill every cached call.
OpenRouter serves an OpenAI-compatible embeddings endpoint, so an ai-embed block may name an llm-openrouter connector. Embedding model ids are vendor-prefixed like every other model here: openai/text-embedding-3-small, google/gemini-embedding-001.
Attachments and generated media
An ai-agent or ai-mapping can send files as well as text, and what a connector accepts is the provider's business. The sets differ, and the differences are real rather than incidental:
| Images | Documents | Audio / video | Returns files | |
|---|---|---|---|---|
llm-anthropic | jpeg, png, gif, webp | pdf, plain text | — | no |
llm-openai | png, jpeg, gif, webp | — | generated images | |
llm-gemini | image/* | pdf, plain text | audio/*, video/* | yes |
llm-openrouter | png, jpeg, webp | — | generated images |
llm-gemini is the only one that reads a voice note or a clip, so a
transcribe-this flow names it. It accepts those by family: image/, audio/
and video/ match anything under them, so a format nobody thought to enumerate —
video/quicktime, image/heif — is still sent.
A type the connector cannot send errors rather than being quietly dropped. A model answering confidently about a file it never received is a failure nothing downstream can detect, so it is not an outcome the runtime allows.
llm-openrouter's row is what this connector can encode, not what will be
accepted. OpenRouter fronts many vendors and what a file is actually taken by is
a property of whichever model the request routes to; a model that refuses it
answers with a provider error.
llm-anthropic returns no files at all — the Messages API has no content block
for one — so responseMedia against it is always empty. Generated media reaches
a flow only when the turn finishes, never progressively
(#517).
Streaming and thinking
All four connectors stream, and an ai-agent with
stream: true reports the model's output to its events path as it is produced.
A streamed turn returns exactly what the same turn would have returned unstreamed
(same text, tool calls, stop reason and usage), so streaming changes when a flow
learns things, never what it learns.
The providers do not agree on what they stream, and the gaps are permanent rather than pending:
| text | thinking | tool arguments | |
|---|---|---|---|
llm-anthropic | fragments | fragments | fragments |
llm-openai | fragments | summary fragments | fragments |
llm-gemini | fragments | fragments | whole, in one event |
llm-openrouter | fragments | fragments | fragments |
Leave reasoning at default unless you want a specific effort. The field is
documented for gpt-5 and o-series models only, and baseURL points this connector
at Azure, proxies and other compatible servers whose model list is not OpenAI's, so
saying nothing is the setting that works everywhere.
OpenAI returns reasoning as a summary rather than verbatim, and only when an
effort other than none is set. The reasoning itself comes back encrypted and is
echoed, unread, on the next turn so a tool loop keeps its train of thought.
Gemini delivers a tool call's arguments complete rather than in pieces; a path
that concatenates tool_input fragments per call still ends up with the same
JSON on all three.
A tool's result goes back to every provider as text, under a single field on
Gemini rather than spread into the function response, because Gemini resolves a
#/... string inside a function response against the response's parts and would
read a tool returning OpenAPI or JSON Schema as dangling $refs. Your tools can
return whatever JSON they like; the model still sees it, as text.
Three things to know before turning thinking on.
Reasoning never reaches the message body. It rides on the assistant turn, not in the answer, so a flow that folds the result into a body cannot publish the model's chain of thought by accident.
Anthropic thinking yields when a request cannot carry it, rather than failing. It
cannot be combined with a forced tool choice, which the API answers with a 400, so
ai-router (which forces one on every call) gets its requests without thinking
while ai-agent is unaffected. Nor with a per-block maxTokens too small to hold
thinkingBudget, which startup validation cannot see because it only knows the
connector's own default.
off and default mean do not ask, not turn off. Every provider here omits
the field entirely for those values, leaving the model on its own default, so on a
model that reasons by default (Gemini 2.5, and many of the reasoning models
OpenRouter fronts) a connector reading thinking: off still reasons on every call.
reasoning: none on llm-openrouter and llm-openai is the one value that
switches it off explicitly; llm-anthropic needs no equivalent because that API
reasons only when asked. llm-gemini currently has no way to express a zero budget.
This matters most where a block sets a small maxTokens. Reasoning is spent before
any answer and counts against the same ceiling, so a reasoning model asked for a
hundred tokens can spend them all thinking and return an empty body, which surfaces
downstream as a parse failure rather than a token error. Until
#344 lands, pair a small
ceiling with a connector that asks for no reasoning.
Example
Declare one or more providers and point an AI block at whichever you like; swapping providers is a one-line change.
service:
name: ai-demo
env:
- name: ANTHROPIC_API_KEY
required: true
connectors:
- name: claude
type: llm-anthropic
settings:
apiKey: ${ANTHROPIC_API_KEY}
model: claude-sonnet-4-6
maxTokens: 2048
flows:
- name: summarize
process:
- type: ai-mapping
settings:
connector: claude # any llm-category connector works here
prompt: Summarize the incoming order into a one-line description.Keep API keys in the environment (${OPENAI_API_KEY}, ${ANTHROPIC_API_KEY}, ${GEMINI_API_KEY}, ${OPENROUTER_API_KEY}), never inline, see environment and config. For what the AI blocks can do with these connectors, start at AI in Octo.