TL;DR — This project builds a production‑grade LoRA inference engine that can serve multiple adapters from a single shared base model, paging weights in and out on demand. It demonstrates real systems skills—memory management, concurrent request handling, and modular architecture—that hiring managers look for.
In this guide you will construct a side project that not only runs inference with multiple LoRA adapters but also showcases your ability to design and implement systems that handle resource constraints, concurrency, and extensibility. The code is written in Python using PyTorch and Hugging Face Transformers, and it can be run on a single GPU or scaled across containers.
Why This Project Stands Out on a CV
- Memory‑efficient weight paging – You’ll implement a custom pager that loads/unloads adapter weights on demand, proving you understand virtual memory concepts in a deep‑learning context.
- Shared base model architecture – By keeping a single instance of the base model and multiplexing adapters, you demonstrate the ability to design resource‑shared services, a pattern common in micro‑service and SaaS back‑ends.
- Concurrent request handling – The included FastAPI server shows you can serve multiple inference requests safely, a skill directly transferable to building production APIs.
- Modular, extensible codebase – The separation of concerns (loader, registry, pager, server) signals you can build maintainable, testable systems.
- Observability & metrics – Integrating Prometheus metrics and structured logging highlights your readiness for production environments.
These competencies map to roles such as ML Engineer, Systems Engineer, or MLOps Engineer, where the ability to ship efficient, scalable inference pipelines is prized.
Architecture Overview
The system is composed of five logical components:
- Base Model Loader – Loads the frozen pre‑trained model (e.g.,
bert-base-uncased) once into GPU memory and exposes aforwardmethod. - Adapter Registry – A thread‑safe dictionary that maps adapter names to their weight files (
.binor.safetensors). - Weight Pager – Implements paging logic: on a cache miss it streams the required adapter weights from disk (or an object store) into a pre‑allocated buffer, then swaps them into the model’s parameter slots.
- Inference Engine – Combines the base model with the currently paged adapter to run inference. It also handles batching and optional quantization.
- API Server – A FastAPI application exposing
/predictendpoints, authenticating requests, and emitting Prometheus metrics.
A simplified textual diagram:
+-------------------+ +-------------------+
| API Server | ---> | Inference Engine |
+-------------------+ +-------------------+
| |
| +-------------+ |
+-->| Weight Pager|<---+
+-------------+
|
v
+-------------------+
| Adapter Registry |
+-------------------+
Building It Step by Step
Step 1 – Set up the environment
python -m venv venv
source venv/bin/activate
pip install torch transformers fastapi uvicorn prometheus-client safetensors
Step 2 – Define the base model loader
# loaders/base.py
import torch
from transformers import AutoModel
class BaseModelLoader:
def __init__(self, model_name: str, device: str = "cuda"):
self.model = AutoModel.from_pretrained(model_name).to(device)
self.model.eval()
self.device = device
def forward(self, input_ids, attention_mask):
with torch.no_grad():
return self.model(input_ids=input_ids,
attention_mask=attention_mask)
Step 3 – Create an adapter registry
# registry/adapter_registry.py
import threading
from pathlib import Path
class AdapterRegistry:
def __init__(self):
self._lock = threading.Lock()
self._adapters = {}
def register(self, name: str, path: Path):
with self._lock:
self._adapters[name] = path
def get(self, name: str) -> Path:
with self._lock:
return self._adapters[name]
Step 4 – Implement the weight pager
# pager/weight_pager.py
import torch
from safetensors.torch import load_file
from registry.adapter_registry import AdapterRegistry
class WeightPager:
def __init__(self, base_loader: BaseModelLoader,
registry: AdapterRegistry,
max_adapters: int = 4):
self.base = base_loader
self.registry = registry
self.max_adapters = max_adapters
self.current = None # name of loaded adapter
def load_adapter(self, name: str):
if name == self.current:
return
path = self.registry.get(name)
# Stream weights from disk
state_dict = load_file(path)
# Replace LoRA parameters in the base model
for key, param in self.base.model.named_parameters():
if key in state_dict:
param.data.copy_(state_dict[key])
self.current = name
Step 5 – Build the inference engine
# engine/inference_engine.py
from loaders.base import BaseModelLoader
from pager.weight_pager import WeightPager
from registry.adapter_registry import AdapterRegistry
class InferenceEngine:
def __init__(self, base_loader: BaseModelLoader,
registry: AdapterRegistry):
self.base_loader = base_loader
self.pager = WeightPager(base_loader, registry)
def predict(self, adapter: str, input_ids, attention_mask):
self.pager.load_adapter(adapter)
return self.base_loader.forward(input_ids, attention_mask)
Step 6 – Expose a FastAPI service
# api/server.py
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from engine.inference_engine import InferenceEngine
from loaders.base import BaseModelLoader
from registry.adapter_registry import AdapterRegistry
import time
app = FastAPI()
# Initialize components
base_loader = BaseModelLoader("bert-base-uncased")
registry = AdapterRegistry()
registry.register("sentiment", Path("adapters/sentiment.safetensors"))
registry.register("ner", Path("adapters/ner.safetensors"))
engine = InferenceEngine(base_loader, registry)
class PredictRequest(BaseModel):
adapter: str
text: str
@app.post("/predict")
async def predict(req: PredictRequest):
# Tokenize (simplified)
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
enc = tokenizer(req.text, return_tensors="pt")
try:
output = engine.predict(req.adapter,
enc["input_ids"],
enc["attention_mask"])
return {"logits": output.logits.tolist()}
except Exception as e:
raise HTTPException(status_code=500, detail=str(e))
Step 7 – Add a CLI for quick testing
# cli/main.py
import argparse
from engine.inference_engine import InferenceEngine
from loaders.base import BaseModelLoader
from registry.adapter_registry import AdapterRegistry
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--adapter", required=True)
parser.add_argument("--text", required=True)
args = parser.parse_args()
base = BaseModelLoader("bert-base-uncased")
reg = AdapterRegistry()
reg.register(args.adapter, Path(f"adapters/{args.adapter}.safetensors"))
engine = InferenceEngine(base, reg)
result = engine.predict(args.adapter, *tokenize(args.text))
print(result)
if __name__ == "__main__":
main()
Running and Testing It
Prepare adapters – Fine‑tune two LoRA adapters (e.g., sentiment, NER) and export them as
.safetensorsfiles in theadapters/directory.Start the API server
uvicorn api.server:app --reload --host 0.0.0.0 --port 8000
- Send a test request
curl -X POST "http://localhost:8000/predict" \
-H "Content-Type: application/json" \
-d '{"adapter":"sentiment","text":"I love this product!"}'
- Run the CLI
python cli/main.py --adapter sentiment --text "Great service!"
- Validate with metrics – Open
http://localhost:8000/metrics(Prometheus) to confirm request counts and latency.
A simple unit test using pytest can verify that the pager swaps adapters correctly:
# tests/test_pager.py
from pager.weight_pager import WeightPager
from loaders.base import BaseModelLoader
from registry.adapter_registry import AdapterRegistry
def test_pager_swaps():
base = BaseModelLoader("bert-base-uncased")
reg = AdapterRegistry()
reg.register("a", Path("adapters/a.safetensors"))
reg.register("b", Path("adapters/b.safetensors"))
pager = WeightPager(base, reg)
pager.load_adapter("a")
assert pager.current == "a"
pager.load_adapter("b")
assert pager.current == "b"
Extending It: Your Roadmap to Senior-Level
- Persistent adapter storage with Redis – Cache adapter weights in Redis to avoid disk I/O on every request, reducing latency and enabling fast scaling.
- Horizontal scaling via Kubernetes – Deploy the API server behind a Kubernetes Deployment with multiple replicas, using a load balancer to distribute traffic.
- Observability stack (Prometheus + Grafana) – Export detailed latency, memory, and error metrics; create dashboards to monitor health in production.
- Fault tolerance with retries and circuit breakers – Wrap external calls (e.g., to a model registry) with retry logic and circuit‑breaker patterns to prevent cascading failures.
- Benchmarking suite with Locust – Simulate realistic traffic patterns, measure throughput and GPU utilization, and identify bottlenecks for optimization.
- Quantization & ONNX export – Convert the base model to ONNX with dynamic quantization, enabling inference on CPU or edge devices and broadening deployment options.
Key Takeaways
- Implemented a weight‑paging mechanism that dynamically swaps LoRA adapters while keeping a single base model in memory.
- Designed a modular architecture (loader, registry, pager, engine, API) that mirrors production service patterns.
- Provided end‑to‑end runnable code, including a FastAPI server, CLI, and test suite, demonstrating full‑stack engineering ability.
- Showcased scalability and observability considerations, positioning the project as a foundation for senior‑level systems work.
Further Reading
- LoRA: Low‑Rank Adaptation of Large Language Models – the original paper that introduced the technique.
- Hugging Face Transformers Documentation – canonical guide for loading and fine‑tuning models.
- FastAPI Official Tutorial – build production‑ready APIs with Python.
- Kubernetes Production Best Practices – scale your inference service horizontally.
- Prometheus Monitoring Guide – instrument your system with metrics.