Diagram of a mini GPT training pipeline showing token embeddings, attention, MLP blocks, and an optimizer loop.

Build a Mini GPT From Scratch in Pure NumPy: A Hiring-Manager-Worthy Project

Build a tiny GPT end-to-end in NumPy to demonstrate ownership of the ML stack: mixed precision, AdamW, cosine LR, and gradient accumulation. Code-first, with a roadmap to production-grade upgrades.

September 6, 2026 · 11 min · 2343 words · martinuke0
Stylized circular ringbuffer with producer and consumer pointers.

Designing Lock-Free Queues: Inside the Linux KFIFO and Ringbuffer Implementations

How the Linux kernel’s KFIFO achieves a wait-free, lock-free FIFO with a single atomic variable. We walk through the ringbuffer math, the memory-ordering subtlety that makes it correct on weakly ordered CPUs, and what you can borrow for your own high-throughput pipelines.

September 6, 2026 · 14 min · 2883 words · martinuke0
Block-tiled attention computation showing memory access patterns across query, key, and value matrices.

Building Flash Attention From Scratch in NumPy: A CV-Grade Side Project

A complete, runnable walkthrough of implementing flash attention from scratch in NumPy, including online softmax, block tiling, causal masking, and correctness checks — designed as a CV-worthy systems project.

September 5, 2026 · 13 min · 2743 words · martinuke0
Diagram of NATS JetStream stream replication across three nodes

Inside NATS JetStream: Building Exactly-Once Delivery with Replicated Logs and Consumer Acks

A deep dive into the architecture behind JetStream’s exactly-once semantics — covering Raft-replicated streams, per-consumer ack tracking, redelivery semantics, and the dedup window that ties it all together.

September 5, 2026 · 8 min · 1637 words · martinuke0
Stylized diagram of a GPU die with attention tile matrices overlaid in cyan and amber.

Optimizing Triton Kernels for Custom Attention Patterns on H100 GPUs

A working engineer’s guide to squeezing the H100 with custom Triton kernels: tile sizing, swizzle patterns, pipeline stages, and the warp-specialization tricks that make non-vanilla attention fly.

September 5, 2026 · 11 min · 2308 words · martinuke0
Feedback