inference
4 posts — newest first.
-
Your cluster has GPUs. That doesn't make it AI-ready.
CNCF's Kubernetes AI Conformance program is a portability standard. Read it as a self-audit instead and it becomes a genuinely useful platform checklist.
-
Round-robin is malpractice for LLM traffic: what the Inference Gateway actually fixes
The Gateway API Inference Extension is GA. Why a normal Kubernetes Service balances model servers badly, and how to tell if you need an InferencePool.
-
GPU utilization is a lying metric: the unit economics of self-hosted inference
You can run a GPU at 100% utilization and waste most of it. The cost model that predicts your inference bill, and the five levers in payback order.
-
Token FinOps: the third budget your agents are spending
Error budgets, context budgets — agents add a third: dollars. Agent tasks burn 5–30× chatbot tokens, and cost-per-token is the wrong metric.