Building a From-Scratch Group-Wise INT4 Quantization Layer for LLaMA Weights
A hands-on guide to implementing group-wise INT4 quantization for LLaMA-style transformers, covering per-group scale calibration, fused dequantization kernels, and benchmarks that demonstrate real ML-systems engineering skill.