Build a RAG Pipeline
Embed text with Gemini, store and search it in Pinecone, and let an agent answer from what it retrieves.
RAG (retrieval-augmented generation) embeds your own content, stores the vectors, and at query time retrieves the most similar before asking a model to answer. This guide builds it with the ai-embed block, the pinecone connector, and an ai-agent that decides for itself when to search. It follows
samples/rag-pipeline.
Three flows, with the middle one shared:
| Flow | Source | What it does |
|---|---|---|
ingest | POST /documents | Embeds a document and upserts it into the index. |
search | none, it is called | Embeds a phrase, returns the closest documents. |
ask | POST /ask | An agent that calls search as a tool, then answers. |
Set up Gemini and Pinecone
This sample needs two real external accounts; there is no local stand-in for a vector index.
Get a Gemini API key from aistudio.google.com/apikey.
Create a Pinecone account and an index at app.pinecone.io
with Dimension 768 (it must match the ai-embed block's dimensions
setting below) and Metric cosine.
Copy the index name and an API key.
Export the credentials and run it:
export GEMINI_API_KEY=...
export PINECONE_API_KEY=...
export PINECONE_INDEX=... # the index name from step 2
task run:sample -- rag-pipelineThe pinecone connector looks its index up at startup and compares the index's
actual dimension against the connector's dimension setting, so a mismatch
fails at boot rather than on the first upsert.
Embed and store: the ingest flow
ai-embed turns body.text into a vector and puts it in vars.vector;
pinecone-upsert writes it into the index, keyed by body.id. Upserting the
same id again overwrites it, so ingestion is idempotent by id.
- name: ingest
source:
connector: api
type: http
settings:
path: /documents
process:
- type: ai-embed
name: embed-document
settings:
connector: gemini
text: body.text
model: gemini-embedding-001
dimensions: 768
resultVar: vector
- type: pinecone-upsert
name: store
settings:
connector: pinecone
vectors: '[{"id": body.id, "values": vars.vector, "metadata": {"text": body.text}}]'ai-embed writes to a variable because pinecone-upsert still needs body.id
and body.text to build the record. pinecone-upsert writes to the body, since
no resultVar names a variable, so the endpoint answers {"upserted": 1} with
no set-payload. Every pinecone-* block follows that rule; see
Results: the body, or a variable.
ai-embed takes an OpenAI- or Gemini-backed connector; Anthropic has no
embeddings API, so an llm-anthropic connector fails at flow build time.
vectors is a single CEL expression evaluating to a list of
{id, values, metadata} objects. Send an array from body and the same block
batches it in one call, chunked automatically if it is large. metadata carries
whatever you want back at search time, here the original text.
curl -s localhost:8080/documents -d '{
"id": "plateau",
"text": "Plateaus in bouldering are usually a training-variety problem, not a strength problem."
}'
# -> {"upserted": 1}Retrieval as a flow of its own
search is the mirror image of ingest (embed, then query instead of upsert)
and it has no source: it is the retrieval step, named once so everything
needing retrieval refers to it.
- name: search
process:
- type: ai-embed
name: embed-query
settings:
connector: gemini
text: body.query
model: gemini-embedding-001
dimensions: 768
resultVar: queryVector
- type: pinecone-query
name: find-similar
settings:
connector: pinecone
vector: vars.queryVector
topK: 3
includeMetadata: trueNo resultVar on the query, so the matches are the flow's body, a list of
{id, score, metadata}. Run retrieval on its own:
octo invoke -config samples/rag-pipeline -flow search -data '{"query": "training plateau"}'
# -> [{"id": "plateau", "score": 0.52, "metadata": {"text": "Plateaus in bouldering are..."}}]The query shares no words with the stored text: the match is on meaning, not keywords.
Answering: retrieval as an agent tool
A fixed chain (embed, query, prompt) fails a question that needs two searches,
or none, so ask hands retrieval to the model as a tool:
- name: ask
source:
connector: api
type: http
settings:
path: /ask
process:
- type: ai-agent
name: answer-question
connector: gemini
maxIterations: 5
prompt: >
Answer the user's question in body.question about their knowledge base.
Call search_documents with a short phrase describing what you need; it
returns the closest documents, each with its id and its text under
metadata. Answer using ONLY the text those documents contain, and
respond with a JSON object {"answer": "...", "sources": ["id", ...]}
listing the ids you used.
guardrail: >
If the documents that come back do not contain the answer, do not guess
from your own knowledge — take the default path instead.
tools:
- name: search_documents
description: >
Search the knowledge base for the documents most similar in meaning
to a phrase. Returns a list of {id, score, metadata}.
inputSchema: |
{
"type": "object",
"required": ["query"],
"properties": { "query": { "type": "string" } }
}
process:
- type: flow-ref
name: retrieve
settings:
flow: search
default:
process:
- type: set-payload
settings:
value: '{"answer": "I could not find that in the knowledge base.", "sources": []}'The tool body is a single flow-ref. An agent tool's
arguments arrive as its branch's message body, so the schema's
{"query": "..."} is what search reads, and its result is its branch's
output body, so the matches go back to the model as-is. A pinecone-query
writing to a variable would need a set-payload here to fish them back out.
curl -s localhost:8080/ask -d '{"question": "I stopped improving. What should I change?"}'
# -> {"answer": "Vary your training rather than chasing strength...", "sources": ["plateau"]}sources tells retrieval from invention, and the guardrail sends a question
the documents cannot answer down the default path instead of into a guess.
text evaluates to a string here, so each call embeds one text and the result
is one vector. Give it a list instead and the result is a list of vectors, in
the same order; see ai-embed.
Multi-tenancy: namespaces
Every pinecone-* block takes a namespace expression (falling back to the
connector's default when empty), so a multi-tenant pipeline routes
body.tenantId to a namespace per request. See
Namespaces.
Testing it without an index
The sample ships three suites, one per flow, and they run in CI with fake
credentials and no network. A configured host skips the startup lookup: given
the index host the connector addresses it directly, so PINECONE_HOST in
samples/.env.test lets these flows be built at all. A refused port proves a
block really calls out: mock the embedder, point PINECONE_HOST at
127.0.0.1:9, and the upsert fails with a dial error naming that address.
- name: the vector really goes to Pinecone, and a broken index fails the flow
input:
data: { id: plateau, text: "Plateaus are a training-variety problem." }
env:
PINECONE_HOST: 127.0.0.1:9
mocks:
ingest.embed-document:
default:
body: { id: plateau, text: "Plateaus are a training-variety problem." }
vars: { vector: [0.1, 0.2, 0.3] }
expect:
error: 'pinecone-upsert: upsert chunk: rpc error: code = Unavailable'The same trick works on the model: the Gemini SDK honours
GOOGLE_GEMINI_BASE_URL, so a case pointing it at the discard port proves
ai-embed and ai-agent drive the model without spending a token. None of it
can prove the model's choice to call the tool; no test may call a model. See
Testing a Flow.
Where to go next
Pinecone connector reference
Every pinecone-* block and setting, including fetch and delete.
ai-embed reference
The embed block's full settings, including batch behaviour.
AI Agents
Tool loops, guardrails, memory, and what the default path is for.
Testing a Flow
Mocks, spies, and the suites that run beside every sample.