NumPy code on a laptop screen illustrating attention matrices

Hands‑On Build Guide: Page‑Transformer Inference with Greedy KV‑Cache Eviction and Paged Attention in Numpy

A step‑by‑step guide to implementing a lightweight page‑transformer with paged attention and KV‑cache eviction, ready to run and extend.

September 16, 2026 · 9 min · 1721 words · martinuke0
A schematic of a mixture-of-experts model routing tokens to experts.

Build a Minimal Mixture-of-Experts Inference Engine in Pure Python

Implement a minimal mixture-of-experts inference engine from scratch in Python, demonstrating systems thinking, custom loss functions, and production-ready patterns.

September 14, 2026 · 8 min · 1616 words · martinuke0
Feedback