TL;DR — This post walks you through building a pure‑Python hybrid RAG engine that layers BM25 sparse retrieval, HNSW vector search, cross‑encoder reranking, and a miniature transformer generator, delivering a tangible, runnable project that showcases systems‑level skill for hiring managers.
Building a portfolio project that double‑dips into information‑retrieval and generation is a rare chance to demonstrate both depth and breadth. Hiring managers on LinkedIn see dozens of “ML portfolio” repos that rely on a single notebook; a hybrid RAG system that you can run, benchmark, and extend signals that you understand how production search pipelines are stitched together, how to trade off latency vs. relevance, and how to ship a Python service end‑to‑end. In this guide you’ll create a pure‑Python hybrid RAG engine that uses BM25 for sparse keyword search, HNSW for dense vector search, a cross‑encoder for reranking, and a tiny GPT‑2‑style transformer for generation. By the end you’ll have a working CLI tool, sample tests, and a clear roadmap to production‑grade upgrades.
Why This Project Stands Out on a CV
A hybrid RAG engine is a concrete showcase of several high‑value skill clusters:
- Information‑retrieval engineering – you design and tune BM25 parameters, invert document frequencies, and balance sparse vs. dense signals.
- Vector‑search architecture – you build an HNSW graph, tune ef/ M‑parameters, and integrate with libraries like FAISS or hnswlib.
- Model‑level reranking – you load a cross‑encoder (e.g.,
ms-marco‑distilroberta‑base‑v2), score query‑document pairs, and demonstrate re‑ranking logic that improves nDCG. - Generative AI plumbing – you load a small transformer (
gpt2), manage tokenizers, and generate conditioned text, illustrating knowledge of inference loops and token‑level control. - Systems thinking – you wire components together with a CLI, add timing metrics, and produce reproducible pipelines.
These capabilities map directly to roles such as search engineer, ML engineer, data platform engineer, and backend engineer who own retrieval‑augmented generation services. Because the entire stack is pure Python, you can demonstrate the project in a interview without needing a separate infra team – you’ve already built, tested, and can run it on a laptop.
Architecture Overview
The system can be visualised as a linear pipeline with three parallel retrieval branches that converge before generation:
[Corpus] ──► (1) BM25 Inverted Index ──►
(2) HNSW Vector Graph ──►
(3) Cross‑Encoder Reranker ──►
▼
[Reranked Docs]
▼
[Generator]
▼
[Generated Answer]
Component breakdown
| Component | Responsibility | Typical Library |
|---|---|---|
| Corpus / Document Store | Holds raw text chunks, provides IDs for downstream steps. | docarray, simple list of str |
| BM25 Sparse Retriever | Token‑based BM25 scoring, fast lookup for keyword queries. | rank_bm25 (pure Python) |
| HNSW Dense Retriever | Approximate nearest‑neighbor search over embeddings. | hnswlib or faiss |
| Embedding Model | Turns passages and queries into vectors (e.g., all-MiniLM-L6-v2). | sentence‑transformers |
| Cross‑Encoder Reranker | Re‑scores top‑k candidates with a deeper model for higher precision. | sentence‑transformers cross‑encoder |
| Generator (tiny transformer) | Produces a fluent answer conditioned on the reranked context. | transformers with gpt2 or distilgpt2 |
| CLI / Orchestrator | Glue code: ingest corpus, build indexes, run query loop, print answer. | argparse, rich for pretty output |
All components are pure Python; no external services are required until you decide to scale.
Building It Step By Step
Below are the core implementation steps. Each step includes a runnable Python snippet with a language tag.
Step 1 – Install dependencies
pip install rank-bm25 sentence-transformers faiss-cpu transformers torch rich
Step 2 – Prepare a small corpus
# corpus.py
corpus = [
"HNSW (Hierarchical Navigable Small World) graphs enable fast approximate nearest‑neighbor search.",
"BM25 is a probabilistic ranking function used in information retrieval to rank documents query relevance.",
"Cross‑encoders compute a relevance score for a query‑document pair, typically yielding higher precision than bi‑encoders.",
"GPT‑2 is a transformer‑based language model that can be fine‑tuned or used zero‑shot for text generation.",
"Python’s `rank_bm25` library provides a simple BM25Okapi implementation out‑of‑the‑box.",
]
Step 3 – Build the BM25 index
# bm25_index.py
from rank_bm25 import BM25Okapi
tokenized_corpus = [doc.split() for doc in corpus]
bm25 = BM25Okapi(tokenized_corpus)
def bm25_search(query, top_k=3):
scores = bm25.get_scores(query.split())
top_idx = sorted(range(len(scores)), key=lambda i: scores[i], reverse=True)[:top_k]
return [(corpus[i], scores[i]) for i in top_idx]
Step 4 – Build the HNSW vector index
# hnsw_index.py
from sentence_transformers import SentenceTransformer
import numpy as np
import hnswlib
embedder = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = embedder.encode(corpus, show_progress_bar=False)
index = hnswlib.Index(space="cosine", dim=embeddings.shape[1])
index.init_index(max_elements=len(corpus), ef_construction=40, M=16)
index.add_items(embeddings)
index.set_ef(50) # search parameter
def hnsw_search(query, top_k=3):
q_emb = embedder.encode([query])[0]
ids, distances = index.search(q_emb, top_k)
return [(corpus[i], 1 - d) for i, d in zip(ids[0], distances[0])]
Step 5 – Load a cross‑encoder reranker
# reranker.py
from sentence_transformers import CrossEncoder
cross_encoder = CrossEncoder("ms-marco-distilroberta-base-v2")
def rerank(query, docs, top_k=3):
pairs = [(query, doc) for doc, _ in docs]
scores = cross_encoder.predict(pairs)
# zip scores with docs and sort
ranked = sorted(zip(docs, scores), key=lambda x: x[1], reverse=True)[:top_k]
return [(doc, score) for (doc, score) in ranked]
Step 6 – Initialise the tiny generator
# generator.py
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "gpt2" # 124 M params, runs quickly on CPU
tokenizer = AutoTokenizer.from_pretrained(model_name)
generator = AutoModelForCausalLM.from_pretrained(model_name)
def generate_answer(query, context, max_new_tokens=150):
prompt = f"Context: {context}\nQuestion: {query}\nAnswer:"
inputs = tokenizer(prompt, return_tensors="pt")
output = generator.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=True, temperature=0.7)
return tokenizer.decode(output[0], skip_special_tokens=True)
Step 7 – Wire everything together in a CLI
# rag_cli.py
import argparse
from .bm25_index import bm25_search
from .hnsw_index import hnsw_search
from .reranker import rerank
from .generator import generate_answer
def retrieve(query, top_k=5):
bm25_hits = bm25_search(query, top_k=top_k)
hnsw_hits = hnsw_search(query, top_k=top_k)
# simple union – you could use a more sophisticated merge strategy
combined = bm25_hits + hnsw_hits
reranked = rerank(query, combined, top_k=3)
return reranked
def main():
parser = argparse.ArgumentParser(description="Pure‑Python hybrid RAG engine")
parser.add_argument("query", type=str, help="Search query")
args = parser.parse_args()
results = retrieve(args.query)
context = " ".join(doc for doc, _ in results)
answer = generate_answer(args.query, context)
print("\n--- Generated Answer ---\n")
print(answer)
if __name__ == "__main__":
main()
Run the CLI with:
python -m rag_cli "What is HNSW good for?"
You should see a short generated answer grounded in the corpus.
Running and Testing It
- Start the CLI –
python -m rag_cli "your query"and verify that the generated text references at least one corpus sentence. - Unit‑test the modules – a minimal test suite (
tests/test_rag.py) can assert thatbm25_searchreturns three results, thathnsw_searchreturns scores between 0 and 1, and thatrerankimproves the top score relative to the raw BM25/ HNSW scores.# tests/test_rag.py import pytest from rag_cli import bm25_search, hnsw_search, rerank def test_bm25_returns_three(): hits = bm25_search("HNSW graph", top_k=3) assert len(hits) == 3 def test_rerank_improves(): docs = [("HNSW enables fast ANN", 0.2), ("BM25 ranks docs", 0.1)] q = "What is HNSW?" ranked = rerank(q, docs, top_k=2) # the reranked list should have a higher first‑item score than the original order assert ranked[0][1] > ranked[1][1] - Benchmark retrieval quality – compute precision@3 and nDCG@3 on a handful of hand‑crafted queries using
evalsfrom theragasrepo (optional, but demonstrates metric‑driven thinking).
If every test passes and the CLI prints sensible answers, you have a functional hybrid RAG engine ready for demonstration.
Extending It: Your Roadmap to Senior‑Level
- Persistent indices with
faissorchroma– save the HNSW graph and BM25 posting lists to disk so the engine starts instantly on restart; matters for CI/CD and repeated demo runs. - Horizontal scaling via FastAPI + uvicorn – expose
/retrieveand/generateendpoints; enables integration into larger services or chatbots and showcases API‑design skill. - Observability with OpenTelemetry – instrument request latency, index build time, and reranker confidence; gives you concrete data to tune ef/ M parameters and demonstrates production‑ready monitoring.
- Fault tolerance with circuit‑breaker pattern – wrap the cross‑encoder call; if the model becomes unavailable, fall back to BM25‑only results, preventing a single point of failure.
- Benchmarking pipeline with RAGAS – automatically compute faithfulness, answer‑ relevancy, and context‑precision scores across a dataset of queries; a concrete metric set that hiring managers love to see.
- Fine‑tuning the generator – replace
gpt2with a domain‑specific checkpoint (e.g.,codegpt) and LoRA‑fine‑tune on your own code snippets; shows you can move from zero‑shot to transfer‑learning workflows.
Each upgrade adds a tangible production concern—latency, scaling, reliability, or measurement—turning the toy into a service you could ship to a modest‑scale deployment.
Key Takeaways
- A pure‑Python hybrid RAG engine demonstrates sparse + dense retrieval, reranking, and generation in a single reproducible artifact.
- The project signals systems‑level competence: index construction, pipeline orchestration, and performance tuning that hiring managers look for in search‑engineer and ML‑engineer roles.
- Using BM25, HNSW, a cross‑encoder, and a tiny transformer gives you concrete experience with the most common building blocks of modern RAG systems.
- The step‑by‑step code is fully runnable; you can clone,
pip install, and query it within minutes. - The roadmap of six upgrades maps directly to production concerns—persistent storage, API serving, observability, fault tolerance, benchmarking, and model fine‑tuning.
Further Reading
- Okapi BM25 – the classic probabilistic ranking function; understand the IDF and free‑parameter tuning.
- HNSW: Hierarchical Navigable Small World Graphs – the original paper that introduced the graph structure used in
hnswlibandfaiss. - Cross‑Encoders for Re‑Ranking – describes how cross‑encoders improve retrieval precision and the trade‑off against latency.
- Sentence‑Transformers Documentation – covers the
all-MiniLM-L6-v2model used for embeddings and theCrossEncoderclass. - FAISS: Facebook AI Similarity Search – alternative vector library; you can swap HNSW for FAISS IVF‑Flat if you need GPU acceleration.
- Transformers Quick‑Start – how to load
gpt2(or any causal LM) and generate text with minimal boilerplate. - RAGAS: RAG Evaluation Suite – a Python library for computing faithfulness, answer‑relevancy, and context‑precision metrics on your own queries.
You now have a complete, runnable hybrid RAG engine, a CV‑worthy story, and a clear path from toy to production‑grade system. Happy building!