A network of connected nodes representing a knowledge graph replacing a flat vector index.

Graphify Replaces the RAG Vector Store: Why Graph-Based Retrieval Is the Next Leap

Graph-based retrieval is displacing the pure vector store in production RAG stacks. Here is how Graphify works, why it outperforms flat embeddings on multi-hop questions, and what it takes to migrate.

September 5, 2026 · 8 min · 1616 words · martinuke0
Diagram of a draft model proposing tokens and a target LLM verifying them in a speculative decoding loop.

Build a Speculative Decoding Sampler: A CV-Worthy Side Project

A complete, runnable Python project that pairs a draft n-gram model with a target LLM to verify multiple tokens per forward pass, plus a roadmap for turning it into production-flavored work.

September 5, 2026 · 12 min · 2509 words · martinuke0
Two stylized neural networks exchanging structured messages instead of a single agent calling many tool APIs.

LLM-to-LLM Communication: Why Hero-to-Hero Beats Hero-to-API

The next leap in agentic AI isn’t a bigger model — it’s how models talk to each other. Here’s why LLM-to-LLM protocols like MCP and Google’s A2A are replacing the brittle hero-to-API pattern.

September 4, 2026 · 9 min · 1895 words · martinuke0
Two pipelines merging — a fast draft model feeding into a verifier that accepts or rejects proposed tokens.

Speculative Decoding: How Production LLM Systems Cut Latency by 2–3x

Speculative decoding trades a small draft model for big latency wins — 2–3x faster token generation with mathematically identical outputs. Here’s how it works and where production systems use it.

September 4, 2026 · 9 min · 1734 words · martinuke0
Diagram showing token positions feeding a transformer decoder with KV cache blocks reused across steps.

KV Caching in Production: How Transformers Trade Memory for Latency

A practitioner’s guide to KV caching: what gets cached, why decode steps are cheap, and how production systems like vLLM and TensorRT-LLM squeeze out more throughput.

September 4, 2026 · 8 min · 1633 words · martinuke0
Feedback