Learning Hub
Lesson #69 of 70
Finance & FinOps AI8 min readAdvanced
Semantic Prompt Caching & Model Cascading (SLM to Frontier)
Slash enterprise inference costs by 40-70% using semantic similarity caching with Redis and automated model cascading from Small Language Models to Frontier LLMs.
Works with:GPTCacheRedis Vector SearchLiteLLM ProxyRouteLLM
Key Takeaways
- Exact-match caching fails when users rephrase identical questions ("How do I reset my password?" vs "Where can I change password?")
- Semantic caching compares vector embeddings of queries against a Redis cache with a cosine similarity threshold (e.g., 0.92+)
- Model Cascading routes 70-80% of routine requests to fast, inexpensive Small Language Models (SLMs like Llama 3 8B or GPT-4o-mini)
- Confidence scoring escalates complex multi-step reasoning queries to Frontier models (Claude 3.5 Sonnet / GPT-4o) only when necessary
The Diagnostic Context
Sending every simple user question to expensive flagship models burns cloud budgets unnecessarily. Combining semantic caching with intelligent model cascading slashes API bills by up to 70% while improving response latency.
The Core Technique
Model Cascading & Router Architecture
CODE / PROMPT
[ User Request ]
|
[ 1. Semantic Cache Check (Redis) ] ---> (Hit: Similarity >= 0.92) ---> [ Return Cached Answer in 15ms ($0.00) ]
| (Miss)
[ 2. Fast Intent Classifier (SLM) ]
|
+---> [ Simple Query (75%) ] ---> [ Route to SLM: Llama-3-8B / GPT-4o-mini ($0.15 / M tokens) ]
|
+---> [ Complex Reasoning (25%) ] ---> [ Route to Frontier: Claude-3.5-Sonnet ($3.00 / M tokens) ]
PYTHON
# Implementation of Intelligent Model Router:
def route_and_execute_query(user_prompt: str, fast_slm, frontier_llm) -> dict:
# 1. Evaluate complexity using lightweight SLM
classifier_prompt = f"Rate query complexity (1 to 5) where 1=simple factual, 5=complex multi-step coding/math. Respond ONLY with digit.\nQuery: {user_prompt}"
rating = int(fast_slm.invoke(classifier_prompt).strip())
if rating <= 2:
# Route to fast/cheap SLM
response = fast_slm.invoke(user_prompt)
return {"model_used": "slm-fast-tier", "cost_tier": "LOW", "answer": response}
else:
# Escalate to frontier model
response = frontier_llm.invoke(user_prompt)
return {"model_used": "frontier-reasoning-tier", "cost_tier": "HIGH", "answer": response}
5-Minute Activation Challenge
Try This Right Now
Implement a basic semantic cache check in Python using cosine similarity: if cosine similarity between a new question and any stored question exceeds 0.90, return the cached answer without calling the LLM!
Tip: Knowledge only becomes capability once you run the prompt yourself.
Comprehension Check
Test Your Instincts (1 Questions)
1