jnachi
Learning Hub
Python Development9 min readIntermediate

Retrieval-Augmented Generation (RAG): Embeddings & Vector Databases

Build robust RAG pipelines in Python: Master semantic text chunking, embedding generation (OpenAI/Ollama), vector indexing, similarity search, and Vector DBs (Chroma, pgvector).

Works with:ChromaDBpgvectorOpenAI Embeddingstiktoken

Key Takeaways

  • RAG grounds LLMs in enterprise private data by retrieving relevant document snippets and injecting them into the prompt context window
  • Text chunking strategies (recursive character splitting with overlap) preserve semantic context without exceeding embedding model token limits
  • Vector embeddings map high-dimensional semantic meaning into numerical vectors compared using Cosine Similarity or Dot Product
  • Vector databases (Chroma, pgvector, Qdrant, Pinecone) index millions of vector embeddings for sub-millisecond approximate nearest neighbor (ANN) search

The Diagnostic Context

LLMs suffer from two fatal flaws: knowledge cutoffs and hallucinations on proprietary corporate data. Retrieval-Augmented Generation (RAG) solves this by fetching verified, relevant passages from an internal vector store and instructing the model to answer strictly based on the retrieved evidence.

The Core Technique

The Complete RAG Architecture Pipeline

DIAGRAM / WORKFLOW
graph TD
    subgraph Ingestion["Ingestion Pipeline (Offline / Background)"]
        Docs["Enterprise PDF/Doc Files"] --> Chunker["Chunker (500 tokens, 10% overlap)"]
        Chunker --> Embedder["Embedding Model (text-embedding-3-small)"]
        Embedder --> VectorDB[("Vector Database (Chroma / pgvector)<br/>- Vector: [0.012, -0.941, ...]<br/>- Metadata: {doc_id, url, page}")]
    end

    subgraph QueryPipeline["Query Pipeline (Real-Time)"]
        UserQ["User: 'What is our refund policy for broken items?'"] --> QEmbed["Embed User Query"]
        QEmbed --> SimSearch["Vector Similarity Search (Top-k = 3)"]
        VectorDB <-->|Cosine Nearest Neighbor| SimSearch
        SimSearch --> AugmentedPrompt["Prompt Assembly:<br/>Context: [Chunk 1, Chunk 2]<br/>User Query: '...'"]
        AugmentedPrompt --> LLM["LLM (GPT-4o / Claude 3.5)"]
        LLM --> Answer["Grounded Fact-Checked Response"]
    end

End-to-End Python RAG Implementation with ChromaDB

PYTHON
import chromadb
from openai import OpenAI

client = OpenAI()
chroma_client = chromadb.Client()
collection = chroma_client.create_collection(name="company_policies")

# 1. Document Ingestion:
documents = [
    "Hardware returns are accepted within 30 days of delivery with original packaging for a full refund.",
    "Software licenses are non-refundable once the activation license key has been generated.",
    "Enterprise customer support operates 24/7 with a guaranteed 15-minute response SLA."
]

# Generate embeddings:
for i, text in enumerate(documents):
    response = client.embeddings.create(input=text, model="text-embedding-3-small")
    embedding = response.data[0].embedding
    collection.add(
        ids=[f"doc_{i}"],
        embeddings=[embedding],
        documents=[text],
        metadatas=[{"category": "policy"}]
    )

# 2. Query & Retrieval:
user_query = "Can I get my money back for an unused software key?"
query_embed = client.embeddings.create(input=user_query, model="text-embedding-3-small").data[0].embedding

results = collection.query(query_embeddings=[query_embed], n_results=2)
retrieved_context = "
---
".join(results["documents"][0])

# 3. Augmented Generation:
system_prompt = f"""You are a helpful customer support agent. Answer the user question strictly using the provided context. If the answer is not in the context, say you do not know.

Context:
{retrieved_context}"""

completion = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[
        {"role": "system", "content": system_prompt},
        {"role": "user", "content": user_query}
    ]
)

print(completion.choices[0].message.content)
5-Minute Activation Challenge

Try This Right Now

Experiment with chunk size and overlap: Ingest a 5-page policy document using 200-character chunks vs 1,000-character chunks. Compare how chunk size impacts the relevance and completeness of retrieved context passages.

Tip: Knowledge only becomes capability once you run the prompt yourself.

Comprehension Check

Test Your Instincts (3 Questions)

1

What is the primary role of an embedding model (e.g., `text-embedding-3-small`) in a RAG system?

2

Why is chunk overlap (e.g., 10–15% overlap) typically used when splitting long documents for RAG ingestion?

3

Which vector distance metric measures the cosine of the angle between two embedding vectors, independent of vector magnitude?