FinOps
3 posts — newest first.
-
GPU utilization is a lying metric: the unit economics of self-hosted inference
You can run a GPU at 100% utilization and waste most of it. The cost model that predicts your inference bill, and the five levers in payback order.
-
AI SRE is officially a category. Buy it like an SRE, not like a CIO.
Gartner's first Market Guide for AI SRE tooling projects 85% adoption by 2029. Six evaluation questions for cutting through the category hype.
-
Token FinOps: the third budget your agents are spending
Error budgets, context budgets — agents add a third: dollars. Agent tasks burn 5–30× chatbot tokens, and cost-per-token is the wrong metric.