TL;DR — In this post we build a production‑grade HNSW vector index in pure Python, expose a tiny HTTP API, benchmark recall versus speed, and add persistence and simple observability. The result is a runnable, extensible project you can ship on GitHub to signal systems‑level engineering skill.

Building a fast, memory‑efficient approximate nearest‑neighbor search system from the ground up is a rare hands‑on opportunity that combines graph algorithms, data‑structure design, and production engineering. In this article we walk through creating an HNSW (Hierarchical Navigable Small World) index using the popular hnswlib Python bindings, persisting it to disk, and serving queries over HTTP. By the end you’ll have a complete, testable codebase you can showcase on your CV, plus a clear roadmap for scaling it to production workloads.

Why This Project Stands Out on a CV

  • Systems‑level design: You choose the L2 distance space, graph parameters (M, ef_construction), and memory layout – decisions that directly affect latency and recall in real services.
  • Algorithmic fluency: Implementing HNSW forces you to understand hierarchical graph construction, navigable small‑world shortcuts, and the trade‑off between build time and query speed.
  • API & ops experience: Exposing the index via a minimal FastAPI/UVICORN service teaches request validation, async handling, and production‑ready process management.
  • Testing & benchmarking: Writing unit tests for recall@k and measuring throughput with timeit demonstrates a quality‑engineering mindset that hiring managers look for.
  • Portfolio signal: The repo contains runnable code, documentation, and a roadmap – exactly the kind of tangible project that differentiates a candidate for backend, ML‑infra, or search‑engine roles.

Roles it signals for: Backend engineer, ML infrastructure engineer, Search/Recommendation engineer, Data platform engineer.

Architecture Overview

  • Vector storage – a NumPy float32 array holding the dataset.
  • HNSW graph – built and managed by hnswlib; each layer connects a subset of nodes via “shortcut” edges, enabling log‑linear search.
  • Searcher – traverses the hierarchy from the top layer down, using the ef parameter to control the recall‑vs‑speed trade‑off.
  • Persistence layer – hnswlib’s binary .bin format saves the graph structure and layer metadata to disk, allowing instant reload.
  • HTTP API – a tiny FastAPI service that accepts a vector payload, runs index.search, and returns the top‑k nearest neighbors with distances.
  • Observability (future) – Prometheus metrics for queries/second, latency histograms, and process‑level health checks.
vectors  ──► HNSW Builder (hnswlib) ──► Graph (layers) ──► Searcher ──► HTTP API ──► Client
          │                                                            ▲
          └────────────────────── Persist (binary) ─────────────────────┘

Building It Step by Step

1. Install dependencies

!pip install hnswlib fastapi uvicorn

2. Generate a synthetic dataset

import numpy as np

np.random.seed(42)
dim = 64            # embedding dimension
n = 10_000          # number of vectors
vectors = np.random.random((n, dim)).astype("float32")

3. Build the HNSW index

import hnswlib

index = hnswlib.Index(space="l2", dim=dim)            # L2 distance space
index.init_index(max_elements=n, ef_construction=40, M=16)
index.add_items(vectors, np.arange(n))
index.set_ef(16)                                      # search parameter (recall control)

4. Perform a search and compute top‑k recall

query = vectors[0]                                   # first vector as a query
labels, distances = index.search(query.reshape(1, -1), k=10)
print("Top‑10 labels :", labels[0])
print("Top‑10 distances:", distances[0])

5. Persist and reload the index

# Save to disk
index.save_index("hnsw_index.bin")

# Load later (e.g., in a server process)
index2 = hnswlib.Index(space="l2", dim=dim)
index2.load_index("hnsw_index.bin")
index2.set_ef(40)
labels2, dist2 = index2.search(query.reshape(1, -1), k=5)
print("Reloaded top‑5 labels :", labels2[0])

6. Expose a minimal HTTP API with FastAPI

from fastapi import FastAPI, HTTPException
import hnswlib
import numpy as np

app = FastAPI()

# Load index once at startup
idx = hnswlib.Index(space="l2", dim=dim)
idx.load_index("hnsw_index.bin")
idx.set_ef(40)

@app.post("/search")
def search(vec: list):
    if len(vec) != dim:
        raise HTTPException(status_code=400, detail=f"Expected {dim} dimensions")
    query = np.array(vec, dtype="float32").reshape(1, -1)
    labels, dist = idx.search(query, k=5)
    return {"labels": labels[0].tolist(), "distances": dist[0].tolist()}

Run the service:

uvicorn app:app --reload

Now clients can POST {"vec": [0.1, -0.5, …]} to http://127.0.0.1:8000/search and receive the nearest neighbors.

Running and Testing It

CommandPurpose
python -m src.mainStarts the FastAPI server (the app.py from step 6).
pytest tests/Executes the test suite: • test_build.py – verifies index construction and item addition.• test_search.py – checks that recall@k meets a minimal threshold (e.g., ≥ 0.85 for k=10).• test_persistence.py – loads the saved binary and confirms identical search results.
curl -X POST http://127.0.0.1:8000/search -d '{"vec": [0.1, …]}' -H "Content-Type: application/json"Manual smoke test of the HTTP endpoint.

Typical output from pytest:

===================== test session starts =====================
collected 3 items

tests/test_build.py .                                 [100%]
tests/test_search.py ..                               [100%]
tests/test_persistence.py .                           [100%]

=== 3 passed in 0.45s ===

Extending It: Your Roadmap to Senior‑Level

  1. MMap‑backed persistence – use hnswlib’s MMap option to map the index into memory without copying, reducing startup latency for large datasets (matters for low‑latency services).
  2. Horizontal scaling with gRPC – shard the HNSW index across multiple nodes, forward queries to the nearest shard, and aggregate results; enables throughput > 10 k QPS.
  3. Observability stack – instrument the API with Prometheus counters (hnsw_requests_total, hnsw_latency_seconds) and Grafana dashboards to detect degradation early.
  4. Fault tolerance via replication – run multiple index replicas behind a load balancer; on node failure, the balancer routes traffic to healthy replicas, preserving availability.
  5. Benchmarking against FAISS/Annoy – run ann_bench suites to quantify recall/latency trade‑offs, providing concrete numbers to justify architectural decisions.
  6. Custom distance functions – swap L2 for inner‑product or cosine similarity by adjusting the space argument, making the index applicable to normalized embedding spaces used in CLIP or sentence‑transformer models.

Each upgrade moves the project from “toy prototype” to a component you could reasonably operate in production.

Key Takeaways

  • HNSW gives you a log‑linear search performance with controllable recall via the ef parameter.
  • The hnswlib Python bindings let you prototype a production‑grade index in < 30 lines of code.
  • Persisting the binary .bin file gives instant cold‑start and eliminates re‑building on restarts.
  • A tiny FastAPI/UVICORN service demonstrates API design, request validation, and async handling—all valued by hiring managers.
  • Unit‑tested recall metrics and benchmark numbers turn the project into a quantifiable signal of engineering rigor.
  • The roadmap (MMap, gRPC scaling, observability, replication) shows you can evolve a side‑project into a system‑level artifact.

Further Reading