A visualization of LLM sampling distributions at different temperature settings, showing probability mass shifting from peaked to flat.

Build a From-Scratch LLM Sampler: Temperature, Top-k/Top-p, and Multinomial Sampling vs. Hugging Face

A hands-on guide to building a from-scratch LLM sampler with temperature scaling, top-k/top-p warping, and multinomial sampling, benchmarked against Hugging Face. Includes runnable code and a roadmap to production.

September 20, 2026 · 14 min · 2888 words · martinuke0
A diagram of a speculative decoding pipeline

Speculative Decoding Engine with Draft-Model Verification and KV-Cache Sharing in Pure Python

This guide walks through building a speculative decoding engine in pure Python, including draft-model verification and KV-cache sharing. It’s ideal for showcasing systems skills to hiring managers.

September 19, 2026 · 8 min · 1589 words · martinuke0
Adaptive nucleus sampler concept

Pure‑Python Adaptive Nucleus Sampler: A Portfolio Project

Implement an adaptive nucleus sampler in Python that adjusts temperature and top‑p from token entropy, demonstrating systems thinking for hiring managers.

September 17, 2026 · 6 min · 1182 words · martinuke0
A minimal GPT training loop diagram

Build a Minimal GPT Training Loop with Rotary Positional Embeddings

Implement a minimal GPT training loop with rotary positional embeddings from scratch to demonstrate real systems skill to hiring managers.

September 17, 2026 · 9 min · 1779 words · martinuke0
A diagram of a context window manager pruning tokens.

Memory‑Efficient Context Window Manager with Attention‑Guided Token Pruning

This project delivers a runnable Python library that compresses LLM context windows by scoring token importance with attention and dropping low‑impact tokens, demonstrating practical systems and ML engineering.

September 17, 2026 · 7 min · 1301 words · martinuke0
Feedback