Hands‑On Build Guide: LoRA‑Based LLM Fine‑Tuning for Your Portfolio CV
A hands‑on guide to building a LoRA‑based LLM fine‑tuning system you can ship to GitHub and discuss in interviews.
A hands‑on guide to building a LoRA‑based LLM fine‑tuning system you can ship to GitHub and discuss in interviews.
Learn to implement a compact LLM inference engine that showcases advanced systems skills, from KV cache reuse to tree-attention speculative decoding.
This tutorial shows how to build a high‑performance top‑k/top‑p token sampler with temperature scaling and per‑token log‑probability tracking using NumPy. It demonstrates real systems skills relevant to ML inference roles.
A practical, runnable pure‑Python LLM inference engine that uses an entropy‑guided sliding‑window KV cache backed by hash‑indexed circular pages – perfect for showcasing systems engineering chops.
A hands‑on guide to building a minimal GPT‑style model from scratch, with runnable code, testing, and extension ideas for a standout portfolio project.