Diagram of grouped-query attention with rotary position embeddings feeding into a tiled Flash-Attention forward pass.

Building Grouped-Query Attention with RoPE and Flash-Attention from Scratch

Step-by-step tutorial for implementing GQA, RoPE, and an online-softmax Flash-Attention forward pass in pure PyTorch — a portfolio project that signals real systems-level ML skill.

September 7, 2026 · 11 min · 2313 words · martinuke0
Block-tiled attention computation showing memory access patterns across query, key, and value matrices.

Building Flash Attention From Scratch in NumPy: A CV-Grade Side Project

A complete, runnable walkthrough of implementing flash attention from scratch in NumPy, including online softmax, block tiling, causal masking, and correctness checks — designed as a CV-worthy systems project.

September 5, 2026 · 13 min · 2743 words · martinuke0
Code editor showing a Python implementation of flash attention with tiling and online softmax in NumPy.

Building Flash Attention from Scratch in Pure NumPy: A CV-Worthy Systems Project

Build flash attention from scratch in pure NumPy using tiling and the online softmax trick, then extend it toward production. A deep, runnable project that signals real systems engineering to hiring managers.

September 4, 2026 · 14 min · 2834 words · martinuke0
Tiled matrix multiplication diagram with softmax rows

Building Flash Attention from Scratch in NumPy: A CV-Worthy Side Project

A working engineer’s guide to implementing Flash Attention from the ground up using tiling and the online softmax trick — runnable NumPy, end-to-end tests, and a roadmap to senior-level extensions.

September 4, 2026 · 10 min · 2047 words · martinuke0

Optimizing Inference Performance Scaling LLM Applications with Quantization and Flash Attention

Table of Contents Introduction Why Inference Performance Matters at Scale Fundamentals of Quantization 3.1 Static vs. Dynamic Quantization 3.2 Post‑Training Quantization (PTQ) Techniques 3.3 Quantization‑Aware Training (QAT) Flash Attention: Reducing Memory Footprint of Self‑Attention 4.1 Algorithmic Overview 4.2 GPU‑Specific Optimizations Putting It All Together: A Practical Pipeline 5.1 Environment Setup 5.2 Quantizing a Hugging Face Model with BitsAndBytes 5.3 Enabling Flash Attention in Transformers 5.4 Benchmarking End‑to‑End Latency and Throughput Scaling Strategies Beyond Quantization & Flash Attention 6.1 Batching & Prefill/Decode Separation 6.2 Tensor Parallelism & Pipeline Parallelism 6.3 Model Sharding on Multi‑GPU Nodes Real‑World Case Studies 7.1 Chatbot Deployment for a Fortune‑500 Customer Service 7.2 Document Retrieval Augmented Generation (RAG) at Scale Best Practices & Common Pitfalls Conclusion Resources Introduction Large language models (LLMs) have moved from research curiosities to production‑grade components powering chatbots, code assistants, and retrieval‑augmented generation pipelines. As model sizes climb into the hundreds of billions of parameters, inference performance becomes a decisive factor for cost, user experience, and environmental impact. Two techniques have risen to the forefront of performance engineering for LLM inference: ...

March 11, 2026 · 11 min · 2197 words · martinuke0
Feedback