Abstract diagram of an agent loop connecting tools, memory, and evaluators

Building and Evaluating Agents: A Production Engineer's Playbook

Most LLM agent demos die in production because teams skip the boring parts: tool contracts, eval harnesses, and observability. This post walks through the architecture, patterns, and evaluation loops that turn a clever prototype into a system you can actually operate.

September 7, 2026 · 13 min · 2563 words · martinuke0
Diagram of an agent loop with planning, tool use, and reflection stages.

Agentic AI in Production: How Autonomous Agents Are Rewiring the Stack

Agentic AI is the shift from one-shot prompting to autonomous loops that plan, call tools, observe results, and recover from failure. This guide walks through the core loop, the production patterns, and the infrastructure that makes agents viable at scale.

September 7, 2026 · 11 min · 2183 words · martinuke0
Diagram of an LLM agent loop connecting perception, planning, tool use, and memory

Agentic AI in 2026: What Stanford's Overview Actually Means for Working Engineers

Stanford’s agentic AI overview offers a clean taxonomy, but production engineers need to know how the four-agent model, tool-use loops, and memory architectures translate into deployed systems — and where they fall apart.

September 7, 2026 · 11 min · 2294 words · martinuke0
Two gears meshing — one large and one small — symbolizing routing work between a heavy frontier model and a lightweight worker model.

The Token Tax: Why Most AI Coding Costs Are Wasted on Work a $0.10 Model Could Do

Most tokens burned by an AI coding agent are spent on I/O and boilerplate, not reasoning. Here’s how to slash your bill by routing grunt work to a cheaper model while keeping your frontier model for the problems that actually need it.

September 6, 2026 · 13 min · 2582 words · martinuke0
Stylized graph diagram showing code symbols as nodes and dependencies as edges.

Context Graphs for Coding Agents: Why Tool-Call Counts Are the New Token Bill

Coding agents spend most of their budget on file and symbol discovery, not on actual reasoning. Here is how precomputed context graphs change that, and why it matters for every team shipping agentic dev tools.

September 6, 2026 · 11 min · 2201 words · martinuke0
Feedback