Hands‑On MoE Router: Top‑k Gating and Load‑Balancing in Pure Python
A hands‑on guide to implementing a Mixture‑of‑Experts router from scratch in pure Python, with top‑k gating and load‑balancing, ready to run and extend.
A Pure‑Python Continuous‑Batching Inference Engine for LLM Requests
A hands‑on guide to building a pure‑Python continuous‑batching inference engine that packs variable‑length LLM requests into GPU‑friendly batches with adaptive timeout and priority scheduling.
Deep Dive into NVIDIA NVLink: Designing High-Bandwidth GPU Topologies for Clusters
A technical guide to NVLink and NVSwitch, covering topology design, bandwidth scaling, and production patterns for GPU clusters.
Deep Dive into BM25+: Tuning Parameters for Large‑Scale Enterprise Search
BM25+ extends classic BM25 with term frequency saturation and document length normalization. This guide walks you through parameter selection, real-world tuning, and production patterns for enterprise search.
Building a PagedAttention KV Cache Manager with Continuous Batching from Scratch
This post walks through building a paged attention KV cache manager with continuous batching, providing runnable Python code and architecture guidance for systems engineers.