Building a Paged KV Cache Allocator with Continuous Batching in Pure Python
This project demonstrates memory management, concurrent request scheduling, and low-level allocator design—all core skills for backend and ML infrastructure roles.
This project demonstrates memory management, concurrent request scheduling, and low-level allocator design—all core skills for backend and ML infrastructure roles.
Implement a paged-attention transformer with LRU KV-cache eviction in pure Python, demonstrating systems engineering proficiency.
Build a production-grade incremental KV cache from scratch in pure Python. This guide covers architecture, implementation, testing, and a roadmap to production-flavored features that signal real systems skill to hiring managers.
A hands-on guide to building a from-scratch LLM decoding sampler implementing temperature, top-k, top-p, and contrastive search in pure Python. Includes architecture, complete code, and a roadmap to senior-level extensions.
Build a pure‑Python speculative decoding engine that uses a draft LLM and parallel acceptance‑reject sampling to speed up text generation. This project highlights real systems engineering skills for your CV.