Abstract visualization of token sampling probabilities

Build a Pure Python LLM Decoder from Scratch: Temperature, Top-k, and Top-p Sampling

A hands-on guide to implementing temperature, top-k, and top-p sampling in pure Python, with runnable code that showcases systems engineering skills.

September 22, 2026 · 10 min · 2026 words · martinuke0
Abstract illustration of a neural network with retrieval arrows

Building a Speculative RAG Decoder with Retrieval-Token Speculation

This tutorial shows how to combine retrieval‑augmented generation with speculative decoding to cut latency while preserving factual grounding. You’ll build a runnable system that demonstrates real‑world performance gains.

September 22, 2026 · 7 min · 1384 words · martinuke0
A diagram of a KV cache with entries being evicted

Build a KV Cache Eviction Engine for LLM Context Windows

Implement a pluggable KV cache eviction engine from scratch, demonstrating systems thinking and production-ready patterns.

September 22, 2026 · 7 min · 1465 words · martinuke0
A code editor displaying Python implementation of paged attention and KV-cache management

Build a Minimal LLM Inference Engine with Paged Attention and Adaptive KV-Cache Eviction in Pure Python

A hands-on guide to building a minimal LLM inference engine with paged attention and adaptive KV-cache eviction in pure Python, designed to demonstrate real systems engineering prowess on your CV.

September 21, 2026 · 13 min · 2752 words · martinuke0
Diagram of a multi-GPU LLM inference pipeline with TensorRT-LLM INT8 quantization

Optimizing TensorRT-LLM: INT8 Quantization Pipelines for Real-Time Inference on Multi-GPU Clusters

A practical guide to INT8 quantization in TensorRT-LLM, covering pipeline architecture, kernel-level optimizations, and multi-GPU deployment strategies for real-time serving.

September 21, 2026 · 9 min · 1816 words · martinuke0
Feedback