Skip to content

Why AI spend behaves differently

  • Usage, not capacity. Model and API costs are often charged per request or per token, so spend follows how much people use a feature, not how many servers you run.
  • Expensive infrastructure. Training and inference on GPUs cost far more per hour than typical compute, and idle capacity is costly.
  • Spread across teams. Experiments start in many places at once, often on separate accounts and tools.
  • Hard to link to value. It is easy to see what a model costs and hard to see what it earns, unless you measure both.

Five practices that keep it under control

  1. Make every cost attributable. Tag and allocate spend to a product, team and use case, including model API usage. Unallocated spend is the first thing to drive down.
  2. Measure unit economics. Track cost per transaction, per customer interaction or per document processed, next to the value each one creates. This is what turns a cost report into a business decision.
  3. Set guardrails in the platform. Budgets and alerts per team, quotas on model usage, and engineering choices that cut cost without hurting quality: routing simple requests to smaller models, caching repeated answers and batching work that is not urgent.
  4. Show cost before deployment. Engineers should see the cost impact of a change before it ships, in the same pipeline that tests it, not when the monthly bill arrives.
  5. Govern it on a fixed rhythm. A monthly review where engineering, finance and product owners look at spend, unit costs and forecasts together, and agree actions.

Who needs to be involved

AI FinOps works when engineering owns cost decisions, finance owns budgets and forecasting, and product owners own the value side. A small FinOps function connects the three, with shared dashboards and agreed definitions of each metric.

A first 90 days

  • Days 1-30: get full visibility. Allocate all cloud and AI spend, find the largest unallocated and idle costs, and agree unit metrics for the main AI features.
  • Days 31-60: act on the obvious waste: idle GPUs, oversized instances, repeated calls that could be cached, and expensive models used for simple tasks.
  • Days 61-90: put guardrails and the monthly governance review in place, and set targets for unit cost alongside the business outcomes each feature supports.

Our SRE and FinOps practice builds this discipline inside client environments, alongside the reliability work that keeps AI services dependable.

Planning a capability centre?

Start with a capability assessment of one function. You keep the findings, with no commitment.