Build a Pure Python LLM Decoder from Scratch: Temperature, Top-k, and Top-p Sampling
A hands-on guide to implementing temperature, top-k, and top-p sampling in pure Python, with runnable code that showcases systems engineering skills.
A hands-on guide to implementing temperature, top-k, and top-p sampling in pure Python, with runnable code that showcases systems engineering skills.
This tutorial shows how to combine retrieval‑augmented generation with speculative decoding to cut latency while preserving factual grounding. You’ll build a runnable system that demonstrates real‑world performance gains.
Implement a pluggable KV cache eviction engine from scratch, demonstrating systems thinking and production-ready patterns.
A hands-on guide to building a minimal LLM inference engine with paged attention and adaptive KV-cache eviction in pure Python, designed to demonstrate real systems engineering prowess on your CV.
A practical guide to INT8 quantization in TensorRT-LLM, covering pipeline architecture, kernel-level optimizations, and multi-GPU deployment strategies for real-time serving.