A visualization of INT4 quantization layers inside a transformer model, showing weight matrices being compressed and dequantized.

Building a From-Scratch Group-Wise INT4 Quantization Layer for LLaMA Weights

A hands-on guide to implementing group-wise INT4 quantization for LLaMA-style transformers, covering per-group scale calibration, fused dequantization kernels, and benchmarks that demonstrate real ML-systems engineering skill.

September 21, 2026 · 12 min · 2553 words · martinuke0
A Python code editor showing GGUF binary parsing logic and tensor dequantization

Building a Pure-Python GGUF Weight Loader with Dequantization for Llama Models

A hands-on guide to building a pure-Python GGUF weight loader with full dequantization support for Llama models — covering binary parsing, quantized tensor reconstruction, and a roadmap to production-grade infrastructure.

September 17, 2026 · 12 min · 2528 words · martinuke0
Diagram of a GGUF file being parsed into numpy tensors with dequantization steps.

Building a GGUF Parser and Dequantizer From Scratch: A Portfolio Project That Signals Real Systems Skill

Parse GGUF, dequantize Q4_0 and Q8_0 tensors, and load Llama weights into numpy. A hands-on guide to a side project that demonstrates real systems engineering on a CV.

September 5, 2026 · 14 min · 2871 words · martinuke0
A laptop screen displaying a GPU shader visualizing quantized tensors.

Implementing WebGPU-Accelerated Quantization: A Deep Dive into High-Performance Local LLaMA Inference

A step‑by‑step guide that shows engineers how to combine WebGPU shaders with LLaMA’s GGML backend to achieve low‑latency, high‑throughput inference on a laptop GPU.

June 1, 2026 · 11 min · 2215 words · martinuke0
A laptop screen displaying a GPU heat map beside a Llama model diagram.

Implementing WebGPU-Accelerated Quantization for Local Llama Inference: A Deep Dive into High-Performance Browser Architectures

A step‑by‑step guide that shows engineers how to run a quantized Llama model inside the browser using WebGPU, with code snippets, performance data, and production‑ready patterns.

May 30, 2026 · 10 min · 2084 words · martinuke0
Feedback