Kubernetes
4 posts — newest first.
-
Round-robin is malpractice for LLM traffic: what the Inference Gateway actually fixes
The Gateway API Inference Extension is GA. Why a normal Kubernetes Service balances model servers badly, and how to tell if you need an InferencePool.
-
GPU utilization is a lying metric: the unit economics of self-hosted inference
You can run a GPU at 100% utilization and waste most of it. The cost model that predicts your inference bill, and the five levers in payback order.
-
Frontier models still fail half your incidents: reading ITBench-AA like an SRE
ITBench-AA put frontier models against 59 real Kubernetes incident diagnoses — all scored below 50%. What the benchmark measures and how to use it.
-
Introduction to YAML!
YAML Ain't Markup Language — a human-readable format for configuration. The syntax, the gotchas, and why it's everywhere in infrastructure.