Octov0.11.7
Guides

Build a RAG Pipeline

Embed text with Gemini, store and search it in Pinecone, and let an agent answer from what it retrieves.

RAG (retrieval-augmented generation) embeds your own content, stores the vectors, and at query time retrieves the most similar before asking a model to answer. This guide builds it with the ai-embed block, the pinecone connector, and an ai-agent that decides for itself when to search. It follows samples/rag-pipeline.

Three flows, with the middle one shared:

FlowSourceWhat it does
ingestPOST /documentsEmbeds a document and upserts it into the index.
searchnone, it is calledEmbeds a phrase, returns the closest documents.
askPOST /askAn agent that calls search as a tool, then answers.

Set up Gemini and Pinecone

This sample needs two real external accounts; there is no local stand-in for a vector index.

Get a Gemini API key from aistudio.google.com/apikey.

Create a Pinecone account and an index at app.pinecone.io with Dimension 768 (it must match the ai-embed block's dimensions setting below) and Metric cosine.

Copy the index name and an API key.

Export the credentials and run it:

export GEMINI_API_KEY=...
export PINECONE_API_KEY=...
export PINECONE_INDEX=...   # the index name from step 2
task run:sample -- rag-pipeline

The pinecone connector looks its index up at startup and compares the index's actual dimension against the connector's dimension setting, so a mismatch fails at boot rather than on the first upsert.

Embed and store: the ingest flow

ai-embed turns body.text into a vector and puts it in vars.vector; pinecone-upsert writes it into the index, keyed by body.id. Upserting the same id again overwrites it, so ingestion is idempotent by id.

samples/rag-pipeline/config.yaml (excerpt)
- name: ingest
  source:
    connector: api
    type: http
    settings:
      path: /documents
  process:
    - type: ai-embed
      name: embed-document
      settings:
        connector: gemini
        text: body.text
        model: gemini-embedding-001
        dimensions: 768
        resultVar: vector
    - type: pinecone-upsert
      name: store
      settings:
        connector: pinecone
        vectors: '[{"id": body.id, "values": vars.vector, "metadata": {"text": body.text}}]'

ai-embed writes to a variable because pinecone-upsert still needs body.id and body.text to build the record. pinecone-upsert writes to the body, since no resultVar names a variable, so the endpoint answers {"upserted": 1} with no set-payload. Every pinecone-* block follows that rule; see Results: the body, or a variable.

ai-embed takes an OpenAI- or Gemini-backed connector; Anthropic has no embeddings API, so an llm-anthropic connector fails at flow build time.

vectors is a single CEL expression evaluating to a list of {id, values, metadata} objects. Send an array from body and the same block batches it in one call, chunked automatically if it is large. metadata carries whatever you want back at search time, here the original text.

curl -s localhost:8080/documents -d '{
  "id": "plateau",
  "text": "Plateaus in bouldering are usually a training-variety problem, not a strength problem."
}'
# -> {"upserted": 1}

Retrieval as a flow of its own

search is the mirror image of ingest (embed, then query instead of upsert) and it has no source: it is the retrieval step, named once so everything needing retrieval refers to it.

samples/rag-pipeline/config.yaml (excerpt)
- name: search
  process:
    - type: ai-embed
      name: embed-query
      settings:
        connector: gemini
        text: body.query
        model: gemini-embedding-001
        dimensions: 768
        resultVar: queryVector
    - type: pinecone-query
      name: find-similar
      settings:
        connector: pinecone
        vector: vars.queryVector
        topK: 3
        includeMetadata: true

No resultVar on the query, so the matches are the flow's body, a list of {id, score, metadata}. Run retrieval on its own:

octo invoke -config samples/rag-pipeline -flow search -data '{"query": "training plateau"}'
# -> [{"id": "plateau", "score": 0.52, "metadata": {"text": "Plateaus in bouldering are..."}}]

The query shares no words with the stored text: the match is on meaning, not keywords.

Answering: retrieval as an agent tool

A fixed chain (embed, query, prompt) fails a question that needs two searches, or none, so ask hands retrieval to the model as a tool:

samples/rag-pipeline/config.yaml (excerpt)
- name: ask
  source:
    connector: api
    type: http
    settings:
      path: /ask
  process:
    - type: ai-agent
      name: answer-question
      connector: gemini
      maxIterations: 5
      prompt: >
        Answer the user's question in body.question about their knowledge base.
        Call search_documents with a short phrase describing what you need; it
        returns the closest documents, each with its id and its text under
        metadata. Answer using ONLY the text those documents contain, and
        respond with a JSON object {"answer": "...", "sources": ["id", ...]}
        listing the ids you used.
      guardrail: >
        If the documents that come back do not contain the answer, do not guess
        from your own knowledge — take the default path instead.
      tools:
        - name: search_documents
          description: >
            Search the knowledge base for the documents most similar in meaning
            to a phrase. Returns a list of {id, score, metadata}.
          inputSchema: |
            {
              "type": "object",
              "required": ["query"],
              "properties": { "query": { "type": "string" } }
            }
          process:
            - type: flow-ref
              name: retrieve
              settings:
                flow: search
      default:
        process:
          - type: set-payload
            settings:
              value: '{"answer": "I could not find that in the knowledge base.", "sources": []}'

The tool body is a single flow-ref. An agent tool's arguments arrive as its branch's message body, so the schema's {"query": "..."} is what search reads, and its result is its branch's output body, so the matches go back to the model as-is. A pinecone-query writing to a variable would need a set-payload here to fish them back out.

curl -s localhost:8080/ask -d '{"question": "I stopped improving. What should I change?"}'
# -> {"answer": "Vary your training rather than chasing strength...", "sources": ["plateau"]}

sources tells retrieval from invention, and the guardrail sends a question the documents cannot answer down the default path instead of into a guess.

text evaluates to a string here, so each call embeds one text and the result is one vector. Give it a list instead and the result is a list of vectors, in the same order; see ai-embed.

Multi-tenancy: namespaces

Every pinecone-* block takes a namespace expression (falling back to the connector's default when empty), so a multi-tenant pipeline routes body.tenantId to a namespace per request. See Namespaces.

Testing it without an index

The sample ships three suites, one per flow, and they run in CI with fake credentials and no network. A configured host skips the startup lookup: given the index host the connector addresses it directly, so PINECONE_HOST in samples/.env.test lets these flows be built at all. A refused port proves a block really calls out: mock the embedder, point PINECONE_HOST at 127.0.0.1:9, and the upsert fails with a dial error naming that address.

samples/rag-pipeline/ingest_test.yaml (excerpt)
- name: the vector really goes to Pinecone, and a broken index fails the flow
  input:
    data: { id: plateau, text: "Plateaus are a training-variety problem." }
  env:
    PINECONE_HOST: 127.0.0.1:9
  mocks:
    ingest.embed-document:
      default:
        body: { id: plateau, text: "Plateaus are a training-variety problem." }
        vars: { vector: [0.1, 0.2, 0.3] }
  expect:
    error: 'pinecone-upsert: upsert chunk: rpc error: code = Unavailable'

The same trick works on the model: the Gemini SDK honours GOOGLE_GEMINI_BASE_URL, so a case pointing it at the discard port proves ai-embed and ai-agent drive the model without spending a token. None of it can prove the model's choice to call the tool; no test may call a model. See Testing a Flow.

Where to go next

On this page