← All work

Case study · AI systems · 04

Forge Physique RAG Assistant

A hybrid-retrieval RAG chatbot that only answers from its own knowledge base — try it live, right on this page.

Built for an online physique-coaching business: dense vector search and BM25 keyword search run in parallel, fused with Reciprocal Rank Fusion, then cross-encoder reranked before an LLM ever sees them. LangGraph adds an optional self-correction loop — retrieve, grade, rewrite, retry — for questions the first pass doesn't answer well.

Built & tested

Stack: LangGraph · Groq · Qdrant Cloud · sentence-transformers · Gradio

Forge Physique RAG Assistant answering a live question with cited sources

Step by step

Five stages, one grounded answer

01 / 05

Cache check

The question is embedded and compared against previously answered questions. A close enough match (cosine similarity above 0.95) returns instantly, with no LLM call at all.

02 / 05

Hybrid retrieve

Two searches run against the knowledge base: a dense vector search in Qdrant (mxbai-embed-large-v1) for meaning, and a BM25 keyword search for exact terms.

03 / 05

Fuse & rerank

Both ranked lists are merged with Reciprocal Rank Fusion, then a cross-encoder reranks the combined candidates and keeps only the top 4 most relevant chunks.

04 / 05

Generate

Groq answers using only those 4 chunks as context — grounded in the knowledge base, never the model's own general knowledge.

05 / 05

Self-correct

Optional "Deep Reasoning" mode: if the retrieved chunks are judged insufficient, the query is rewritten and retrieval retries once before answering.

Try it live

Running right now on HF Spaces

Ask it anything about programs, pricing, or coaching policy — it only knows what's in its own knowledge base.

Why it stays on-topic

A support chatbot that makes things up is worse than no chatbot. So every answer is grounded only in the retrieved chunks — the model is never asked to answer from general knowledge, only from what hybrid search actually found.

Dense search alone misses exact terms; keyword search alone misses paraphrased questions. Running both and fusing the results catches what either one would drop on its own.

The semantic cache isn't just a speed trick: a repeated or near-duplicate question gets the same grounded answer every time, instead of a fresh generation that could drift.

Built with

Tech used on this project

PythonLangGraphGroq APIQdrant Cloud sentence-transformersCross-Encoder RerankingBM25GradioHF Spaces

Want a RAG chatbot grounded in your own content?

Open to freelance and contract work on agentic AI systems — RAG pipelines, hybrid search, or something custom. If you can describe the problem, I can tell you honestly whether it is a good fit.

Project
Forge Physique RAG Assistant
Status
Built & tested
Role
Design & engineering
Contact
affanned399@gmail.com