cost-optimization
3 posts — newest first.
-
GPU utilization is a lying metric: the unit economics of self-hosted inference
You can run a GPU at 100% utilization and waste most of it. The cost model that predicts your inference bill, and the five levers in payback order.
-
Most of your AI platform's traffic doesn't need a frontier model
Small models got good enough in 2026. But the win isn't using them — it's making model choice a platform tier with a router, eval gate, and demotion path.
-
Token FinOps: the third budget your agents are spending
Error budgets, context budgets — agents add a third: dollars. Agent tasks burn 5–30× chatbot tokens, and cost-per-token is the wrong metric.