From-Scratch Mini GPT Training Loop with Gradient Checkpointing, Mixed Precision, and a Custom CUDA‑Style Optimizer
A hands‑on guide to training a tiny GPT model end‑to‑end with gradient checkpointing, FP16 mixed precision, and a hand‑rolled optimizer, perfect for a CV‑worthy portfolio.
User Safety in Production Systems: Safe-by-Design Patterns for Engineers
How to build systems that protect users without sacrificing velocity, using proven architecture patterns and concrete tooling.
Dynamic Sliding-Window Context Manager for LLM Prompt Histories
A hands‑on guide to building a production‑ready context manager that automatically trims prompt histories using attention‑weighted saliency, keeping your LLM prompts within any context window.
Architecting Real-time Data Pipelines with Materialize: A Streaming SQL Deep Dive
Explore how Materialize’s streaming SQL engine transforms real-time data pipelines, from Kafka ingestion to downstream consumption with sub-second latency.
Implementing Consistent Bulk Operations Across Sharded MongoDB Clusters: A Deep Dive into Two-Phase Commits
When sharding splits your data across nodes, bulk operations can no longer rely on single-document atomicity. This post walks through implementing consistent two-phase commits in MongoDB clusters to preserve integrity at scale.