Diagram of paged KV cache blocks shared across sequence slots in an LLM inference engine.

Build a Paged KV Cache With Continuous Batching and Prefix Sharing: A From-Scratch Guide

A working engineer’s walkthrough of building a paged KV cache with continuous batching and prefix sharing from scratch — the same ideas vLLM ships in production. Includes runnable Python, an architecture breakdown, tests, and a concrete extension roadmap.

September 7, 2026 · 12 min · 2390 words · martinuke0
Diagram of a transformer training loop showing activations dropped at checkpoint boundaries and recomputed during backward pass.

Building a Gradient Checkpointing Memory Scheduler From Scratch: A Portfolio Project That Actually Signals ML Systems Skill

A practical build guide for a portfolio-grade gradient checkpointing scheduler that trades compute for memory in mini GPT training. Includes runnable PyTorch code, a scheduler design, tests, and a roadmap to production-flavored upgrades.

September 7, 2026 · 15 min · 3026 words · martinuke0
Diagram of a GGUF file layout streaming into memory-mapped safetensors buffers.

Building a GGUF Parser from Scratch: A Portfolio Project That Signals Real Systems Skill

Walk through designing and implementing a GGUF parser that streams tensor shards into memory-mapped safetensors-compatible buffers — a CV-grade systems project with concrete code, tests, and a roadmap to senior-level upgrades.

September 7, 2026 · 14 min · 2820 words · martinuke0
Diagram of a speculative decoding loop with an n-gram draft model feeding into a target verification step.

Building a Speculative Decoding Engine From Scratch: A Portfolio Project That Actually Signals Systems Skill

Build a runnable speculative decoding engine in Python with an n-gram draft model and target-model verification loop — a portfolio project that demonstrates LLM inference, KV-cache reasoning, and acceptance sampling.

September 7, 2026 · 14 min · 2845 words · martinuke0
Diagram of paged-attention KV cache blocks mapped across sequences

Build a Paged-Attention KV Cache Manager: A Hands-On Project That Actually Impresses Hiring Managers

Ship a runnable paged-attention KV cache manager that demonstrates real LLM systems skill: block tables, prefix-sharing across requests, and eviction policies that actually move the needle on memory.

September 7, 2026 · 14 min · 2851 words · martinuke0
Feedback