Your AI Bill Is a Management Accounting Problem

Five large engineering organisations compared notes on runaway AI coding spend. The first tool finance reaches for came third on their list, and they call it a last resort.
In August, Databricks published the first credible multi-company account of enterprise AI coding costs, written with review input from infrastructure leaders at Stripe, Coinbase, Uber and Ramp. The opening admission is the useful part: agentic coding measurably improved every velocity metric they track, and nearly every company deploying at scale hit the same wall, costs compounding on a curve that would eventually overtake revenue. The productivity case and the cost case are on the same graph, and the second one bends faster.
Most companies respond by rationing access. That protects the budget and kills the transformation. What the five converged on instead is a sequence, and the sequence is the finding.
The Biggest Lever Is a Buying Decision
Frontier labs compete on peak intelligence. Enterprises at scale are buying something else: the cheapest model that clears the quality bar for ordinary work. Models clearing that bar at better prices arrive nearly weekly, so the largest single lever is moving spend down to them. The obstacle is measurement. Public benchmarks track poorly to real coding work, so each of these companies built internal evaluations on their own codebases.
The detail that proves the discipline: the evaluations frequently return negative results, and the companies act on them. Stripe measured a newer flagship model against its predecessor, found no quality gain at a higher price, and declined to roll it out. Databricks reached the same conclusion on a different pair. Two of the most sophisticated AI buyers in the market declined an upgrade because their own numbers said worse value. If your evaluation has never told you no, it is a procurement formality, not an evaluation.
Route by Task, Not by Taste
The second lever removes the choice from the user: route each request or task to the cheapest model capable of it. Renaming a component goes to a cheap model; a latency redesign goes to an expensive one. Databricks reports its router cuts average task cost by more than 30 per cent while roughly matching the quality of the most expensive model in the set.
Why Budgets Come Last

Hard monthly caps came third, and every company described them as a last resort, for one reason worth restating to any finance team: some of the highest spenders are the people getting the largest gains. A flat cap punishes exactly the behaviour the adoption programme exists to produce. The working alternative is a ladder: real-time spend visibility for the individual, self-clearing gates that notify at thresholds, then downshifting to a cheaper model rather than cutting access, and suspension only as the opening of a conversation. Friction rises with spend. Capability never reaches zero.
Govern the unit cost and the routing, and the total governs itself. Cap the total and you tax your best adopters first.
One caution the source does not carry far enough: its savings are measured on the input side, and its quality claims are asserted rather than measured. Before copying any lever, decide what output measure will tell you the cheaper model is not silently costing you rework. Savings without a yield measure is half a ledger.
Part of the Business Process Intelligence series from KG Consultancy.
Strategy and technology are the same decision. Over 15 years in fintech (CTOS, D&B), prop-tech (PropertyGuru DataSense), and digital startups, I have built frameworks that help founders and executives make both moves at once. Based in Kuala Lumpur.
Working on a 0→1 product?
I help founders and operators go from idea to validated product. Let's talk about yours.
Get in touch →