Retrieval-Augmented Generation (RAG): Embeddings & Vector Databases
Build robust RAG pipelines in Python: Master semantic text chunking, embedding generation (OpenAI/Ollama), vector indexing, similarity search, and Vector DBs (Chroma, pgvector).
Key Takeaways
- RAG grounds LLMs in enterprise private data by retrieving relevant document snippets and injecting them into the prompt context window
- Text chunking strategies (recursive character splitting with overlap) preserve semantic context without exceeding embedding model token limits
- Vector embeddings map high-dimensional semantic meaning into numerical vectors compared using Cosine Similarity or Dot Product
- Vector databases (Chroma, pgvector, Qdrant, Pinecone) index millions of vector embeddings for sub-millisecond approximate nearest neighbor (ANN) search
The Diagnostic Context
LLMs suffer from two fatal flaws: knowledge cutoffs and hallucinations on proprietary corporate data. Retrieval-Augmented Generation (RAG) solves this by fetching verified, relevant passages from an internal vector store and instructing the model to answer strictly based on the retrieved evidence.
The Core Technique
The Complete RAG Architecture Pipeline
graph TD
subgraph Ingestion["Ingestion Pipeline (Offline / Background)"]
Docs["Enterprise PDF/Doc Files"] --> Chunker["Chunker (500 tokens, 10% overlap)"]
Chunker --> Embedder["Embedding Model (text-embedding-3-small)"]
Embedder --> VectorDB[("Vector Database (Chroma / pgvector)<br/>- Vector: [0.012, -0.941, ...]<br/>- Metadata: {doc_id, url, page}")]
end
subgraph QueryPipeline["Query Pipeline (Real-Time)"]
UserQ["User: 'What is our refund policy for broken items?'"] --> QEmbed["Embed User Query"]
QEmbed --> SimSearch["Vector Similarity Search (Top-k = 3)"]
VectorDB <-->|Cosine Nearest Neighbor| SimSearch
SimSearch --> AugmentedPrompt["Prompt Assembly:<br/>Context: [Chunk 1, Chunk 2]<br/>User Query: '...'"]
AugmentedPrompt --> LLM["LLM (GPT-4o / Claude 3.5)"]
LLM --> Answer["Grounded Fact-Checked Response"]
end
End-to-End Python RAG Implementation with ChromaDB
import chromadb
from openai import OpenAI
client = OpenAI()
chroma_client = chromadb.Client()
collection = chroma_client.create_collection(name="company_policies")
# 1. Document Ingestion:
documents = [
"Hardware returns are accepted within 30 days of delivery with original packaging for a full refund.",
"Software licenses are non-refundable once the activation license key has been generated.",
"Enterprise customer support operates 24/7 with a guaranteed 15-minute response SLA."
]
# Generate embeddings:
for i, text in enumerate(documents):
response = client.embeddings.create(input=text, model="text-embedding-3-small")
embedding = response.data[0].embedding
collection.add(
ids=[f"doc_{i}"],
embeddings=[embedding],
documents=[text],
metadatas=[{"category": "policy"}]
)
# 2. Query & Retrieval:
user_query = "Can I get my money back for an unused software key?"
query_embed = client.embeddings.create(input=user_query, model="text-embedding-3-small").data[0].embedding
results = collection.query(query_embeddings=[query_embed], n_results=2)
retrieved_context = "
---
".join(results["documents"][0])
# 3. Augmented Generation:
system_prompt = f"""You are a helpful customer support agent. Answer the user question strictly using the provided context. If the answer is not in the context, say you do not know.
Context:
{retrieved_context}"""
completion = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_query}
]
)
print(completion.choices[0].message.content)
Try This Right Now
Experiment with chunk size and overlap: Ingest a 5-page policy document using 200-character chunks vs 1,000-character chunks. Compare how chunk size impacts the relevance and completeness of retrieved context passages.
Tip: Knowledge only becomes capability once you run the prompt yourself.