Diagram of a mixed-precision training loop with gradient checkpointing for a mini GPT.

Build a Mini GPT From Scratch: A Mixed-Precision Training Loop That Actually Impresses

Ship a portfolio project that signals real systems skill: a custom autograd engine, mixed-precision training, gradient checkpointing, and AdamW, all wired into a mini GPT you can actually run.

September 7, 2026 · 12 min · 2430 words · martinuke0
Stylized diagram of a rotary positional embedding applied to a query and key tensor on a GPU

Building a Rotary Positional Embedding Kernel from Scratch: A Portfolio Project That Signals Real Systems Skill

Build a real rotary positional embedding kernel in PyTorch with a custom autograd node, fused CUDA-style kernels, and a tiny test harness. The project you ship on a CV to signal GPU, systems, and ML-engineering depth.

September 6, 2026 · 11 min · 2218 words · martinuke0
Block-tiled attention computation showing memory access patterns across query, key, and value matrices.

Building Flash Attention From Scratch in NumPy: A CV-Grade Side Project

A complete, runnable walkthrough of implementing flash attention from scratch in NumPy, including online softmax, block tiling, causal masking, and correctness checks — designed as a CV-worthy systems project.

September 5, 2026 · 13 min · 2743 words · martinuke0
Diagram of mmap-backed tensor parsing pipeline reading GGUF and safetensors files.

Build a GGUF Parser and Safetensors Loader: A From-Scratch Mmap-Backed Weight Inspector

Build a from-scratch GGUF parser and safetensors loader that memory-maps tensor data straight into numpy, with integrity checks you can drop on a CV.

September 5, 2026 · 13 min · 2734 words · martinuke0
Diagram of tiled attention matrix blocks flowing through online softmax accumulation.

Building a FlashAttention-2 Forward Kernel in Pure NumPy

TL;DR — FlashAttention-2 is an IO-aware attention algorithm that fuses the softmax and matrix multiply into tiled CUDA kernels. We rebuild the forward pass from scratch in pure NumPy: blocking the QKV matrices into tiles, running a streaming softmax per block, and rescaling outputs online. The result is a runnable, testable side project that demonstrates memory-hierarchy thinking, numerical stability, and GPU-style programming patterns — exactly what hiring managers look for on a CV. ...

September 5, 2026 · 13 min · 2609 words · martinuke0
Feedback