01 / 05
Cache check
The question is embedded and compared against previously answered questions. A close enough match (cosine similarity above 0.95) returns instantly, with no LLM call at all.
Case study · AI systems · 04
A hybrid-retrieval RAG chatbot that only answers from its own knowledge base — try it live, right on this page.
Built for an online physique-coaching business: dense vector search and BM25 keyword search run in parallel, fused with Reciprocal Rank Fusion, then cross-encoder reranked before an LLM ever sees them. LangGraph adds an optional self-correction loop — retrieve, grade, rewrite, retry — for questions the first pass doesn't answer well.
Five stages, one grounded answer
01 / 05
The question is embedded and compared against previously answered questions. A close enough match (cosine similarity above 0.95) returns instantly, with no LLM call at all.
02 / 05
Two searches run against the knowledge base: a dense vector search in Qdrant (mxbai-embed-large-v1) for meaning, and a BM25 keyword search for exact terms.
03 / 05
Both ranked lists are merged with Reciprocal Rank Fusion, then a cross-encoder reranks the combined candidates and keeps only the top 4 most relevant chunks.
04 / 05
Groq answers using only those 4 chunks as context — grounded in the knowledge base, never the model's own general knowledge.
05 / 05
Optional "Deep Reasoning" mode: if the retrieved chunks are judged insufficient, the query is rewritten and retrieval retries once before answering.
Running right now on HF Spaces
A support chatbot that makes things up is worse than no chatbot. So every answer is grounded only in the retrieved chunks — the model is never asked to answer from general knowledge, only from what hybrid search actually found.
Dense search alone misses exact terms; keyword search alone misses paraphrased questions. Running both and fusing the results catches what either one would drop on its own.
The semantic cache isn't just a speed trick: a repeated or near-duplicate question gets the same grounded answer every time, instead of a fresh generation that could drift.
Tech used on this project
Open to freelance and contract work on agentic AI systems — RAG pipelines, hybrid search, or something custom. If you can describe the problem, I can tell you honestly whether it is a good fit.