Optimizing Small Language Models for Local Edge Inference: Quantization, Pruning, and Runtime Tuning
A field guide to running small language models locally on CPUs, NPUs, and modest GPUs. Covers INT4/INT8 quantization, structured and unstructured pruning, KV-cache tuning, batching, and how to measure what actually matters.