The Economics of Intelligence, Part 2: The True Cost of Intelligence

The cheapest model can produce the most expensive outcome.A team I worked with switched their document processing agent to a cheaper model. The inference bill fell by about sixty percent, and everyone was pleased. Three months later, the true cost of that workflow had gone up, not down.The cheaper…

The cheapest model can produce the most expensive outcome.A team I worked with switched their document processing agent to a cheaper model. The inference bill fell by about sixty percent, and everyone was pleased. Three months later, the true cost of that workflow had gone up, not down.The cheaper model was good, but not quite good enough. Its answers needed a second look more often, so the compliance team started reading every one and correcting more of them. The saving on the invoice was real. It was just smaller than the extra human hours it created, and human hours are the most expensive line in the whole system. They had optimised the number they could see and quietly inflated the one they could not.In Part 1, I argued that intelligence has a meter now, and that the number on it, total tokens consumed, tells you almost nothing about value. This piece is about what to measure instead. It comes down to two moves. Fix the denominator. Then fix the numerator.Fix the denominator: measure work, not consumptionThe first move is to stop dividing by the wrong thing.Most AI reporting measures cost per token, or tokens per user, or tokens per agent. All of these count consumption. None of them count work. The useful denominator is the outcome: cost per resolved claim, tokens per contract reviewed, cost per customer issue closed, cost per decision made.OpenAI made the same point in its recent scorecard for the AI age, framing the goal as useful intelligence per dollar. The instinct is right. A million tokens consumed is not an achievement. It could be extraordinary productivity, or extraordinary waste, and the raw number will never tell you which.To measure work, you first have to define what “done” means for the workflow. A customer issue actually resolved. Code that passes its tests. A contract reviewed correctly and on time. That definition becomes the denominator everything else hangs on, and most organisations have never written it down. Until you do, you are measuring effort and calling it output.Fix the numerator: the token is only one lineThe second move is to stop pretending the token price is the cost.The true cost of an AI outcome is a stack, and inference is only the bottom of it:Model, plus tools and external APIs, plus the compute and orchestration around it, plus the context you loaded, plus the retries when it failed, plus the human review, plus the verification and rework when the answer was wrong.Add those up and the picture inverts. This is why the team in my opening example lost money by saving money. They cut the first line and inflated the last one. A model that is cheaper per call but wrong more often does not reduce cost. It relocates it from the inference bill, where finance can see it, to human time, where finance cannot.The verification taxThat relocated cost has a name I have used before: the Verification Tax. It is the cost of a human re-checking work the machine already did, not because the check adds value, but because trust in the output was never established.The Verification Tax is where dependability becomes economics. OpenAI’s scorecard suggests sorting AI outputs three ways, and it is a good habit: ready to use, needs correction, needs escalation. Those three states are not a quality metric sitting outside the economics.They are the economics. A workflow where most outputs are ready to use is cheap to run even on an expensive model. A workflow where most outputs need correction is expensive to run even on a cheap one.Put concrete numbers on it. Agent A costs ten pence of inference per attempt, but its answers pull a human in to fix them, at three pounds a time. Agent B costs forty pence of inference, four times more, but its answers need only a twenty-pence glance. Agent A looks four times cheaper on the token bill. Agent B is roughly seven times cheaper in reality. If you optimise on the token price, you choose Agent A, and you are worse off every single day.Cost per successful outcomePut the two moves together, and you get the number that actually matters:Total intelligence cost, the full stack, divided by successful outcomes, the ones that met the definition of done.Cost per successful outcome does something cost per token can never do. It counts the failures. A failed agent run still burns tokens, still calls tools, still consumes context, and then produces nothing you can use. Cost per token treats that run as spend. Cost per successful outcome treats it as what it is: waste that also raises the price of every good outcome around it.So the sentence to carry out of this piece is short. Lowest inference cost does not mean lowest outcome cost. Very often it means the opposite.One thing to try this weekPick one workflow where AI is doing real volume. Do not look at its token bill. Instead, work out, even roughly, the cost per successful outcome: the model, the tools, the retries, and crucially the human time spent reviewing and correcting, all divided by the number of outcomes that actually met the bar.Then compare that to what the same workflow would cost with a stronger model that needed less checking. For a surprising number of workflows, the more expensive model is the cheaper answer. That single comparison has changed more allocation decisions, in my experience, than any dashboard of token consumption ever has.What cost per outcome still cannot tell youThere is one more limit, and it is the reason there is a Part 3.Cost per successful outcome prices a workflow. It tells you what one type of work costs to do well. But it still cannot tell you where your intelligence actually went across the whole enterprise, which agent or business unit or customer consumed it, why, and whether the value it created was worth more than the value it consumed. You can price the work. You cannot yet trace a pound of intelligence to the outcome it helped create, or decide where the next pound should go.Pricing is not attribution, and attribution is not allocation. That is the final move, and the heart of the whole series. In Part 3 we build the discipline that does it, and give it its name: Intelligence Accounting.Frequently asked questionsWhy is cost per token the wrong measure? Because it counts consumption, not work, and it ignores everything around the token. A cheaper model can require more retries, more human correction, and more rework, so a lower token price can produce a higher total cost per successful outcome. Measure the outcome, not the input.What is the Verification Tax? The cost of a human re-checking AI output they rarely change, because trust was never established. It is where reliability becomes economics: an unreliable but cheap model quietly moves cost off the inference bill and onto human time, which is far more expensive.This is Part 2 of a three-part series, The Economics of Intelligence. Part 3, From Token Accounting to Intelligence Accounting, follows on Friday. Part 1, Intelligence Has a Meter Now, is here. Prasad Prabhakaran is the author of The AI-Native Enterprise. Early reader access at ainativeenterprise.xyz.This story is published on Generative AI. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories.Subscribe to our newsletter and YouTube channel to stay updated with the latest news and updates on generative AI. Let’s shape the future of AI together!The Economics of Intelligence, Part 2: The True Cost of Intelligence was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

Source: Generative AI Pub — Published — Category: Image AI

🔗 Read full article on Generative AI Pub →