Abstract visualization of AI infrastructure layers blending orchestration, networking, observability, and cost optimization.

From Terraform to Token Economics: The Six Skills That Became One Discipline

AI infrastructure is no longer a subfield of MLOps — it is the convergence of six traditional disciplines. Here is how Terraform, Kubernetes, GPU scheduling, inference routing, SRE, and FinOps are collapsing into a single operating model.

September 5, 2026 · 12 min · 2393 words · martinuke0
Stylized dashboard showing GPU utilization bars alongside cost and revenue metrics.

GPU Utilization Is Becoming a Business Metric

GPUs are no longer a hidden infrastructure cost. They are a tracked business metric, tied to revenue, margins, and team accountability. This post unpacks why utilization moved from engineering dashboards to board decks, and what it means for the way you build, schedule, and bill ML workloads.

September 5, 2026 · 9 min · 1800 words · martinuke0
Layered diagram of an AI inference request flowing from a user app through gateway, inference, GPU, and network layers.

From Prompt to Packet: Tracing an AI Request Through the Modern Inference Stack

We follow a single agent request from the user prompt down through the AI gateway, inference server, GPU/Kubernetes scheduler, and east-west networking — then turn around and look at how observability, security, and FinOps wrap around the whole thing.

September 5, 2026 · 12 min · 2445 words · martinuke0
Feedback