// TODO: I’m martinuke0

Welcome to my corner of the internet. This website is a personal blog which I use as a platform to document my learning journey and showcase it for the world to see.
A sleek Python code editor with neural network visualisation

Hands‑On MoE Router: Top‑k Gating and Load‑Balancing in Pure Python

A hands‑on guide to implementing a Mixture‑of‑Experts router from scratch in pure Python, with top‑k gating and load‑balancing, ready to run and extend.

October 1, 2026 · 9 min · 1864 words · martinuke0
Python code on a GPU dashboard

A Pure‑Python Continuous‑Batching Inference Engine for LLM Requests

A hands‑on guide to building a pure‑Python continuous‑batching inference engine that packs variable‑length LLM requests into GPU‑friendly batches with adaptive timeout and priority scheduling.

October 1, 2026 · 9 min · 1895 words · martinuke0
NVIDIA NVLink GPU cluster topology

Deep Dive into NVIDIA NVLink: Designing High-Bandwidth GPU Topologies for Clusters

A technical guide to NVLink and NVSwitch, covering topology design, bandwidth scaling, and production patterns for GPU clusters.

October 1, 2026 · 6 min · 1197 words · martinuke0
BM25+ formula visualized on a search interface

Deep Dive into BM25+: Tuning Parameters for Large‑Scale Enterprise Search

BM25+ extends classic BM25 with term frequency saturation and document length normalization. This guide walks you through parameter selection, real-world tuning, and production patterns for enterprise search.

October 1, 2026 · 10 min · 2089 words · martinuke0
A diagram of a paged attention KV cache manager

Building a PagedAttention KV Cache Manager with Continuous Batching from Scratch

This post walks through building a paged attention KV cache manager with continuous batching, providing runnable Python code and architecture guidance for systems engineers.

October 1, 2026 · 9 min · 1828 words · martinuke0
Feedback