For years, if someone in financial technology mentioned βtokenomics,β there was a good chance the conversation was about cryptocurrency: supply, distribution, scarcity and incentive design around blockchain tokens.
That definition is still relevant, but a second version of tokenomics is rapidly becoming more important to enterprise technology leaders, and particularly to banks: the economics of the tokens consumed by AI.
I spend much of my time today talking with large financial institutions about moving AI beyond experimentation and into production, across conversational AI, voice, servicing, operational workflows and increasingly agentic systems.Β
One question is coming up much more often: what does all of this really cost when you turn it on at scale?
For the first few years of enterprise generative AI, most of the conversation centered around what the technology could do. Now the conversation is shifting toward what happens when those capabilities are used millions of times, across thousands of employees, customers, workflows and agents.
In other words, we are entering the part of the AI conversation where Finance walks into the room.
The strange economics of cheaper intelligence
The interesting thing about AI tokenomics is that the underlying unit cost is moving in the right direction.
The Washington Post recently reported that the blended price per million AI tokens had declined roughly 50% year over year. Competition between model providers continues to lower inference costs, and enterprises are becoming much more deliberate about routing work toward the least expensive model capable of completing a task.
That sounds like it should mean AI is getting cheaper, but it does not necessarily work that way.
Accenture cites a Goldman Sachs projection that global token consumption could reach roughly 120 quadrillion tokens per month by 2030, around 24 times today's level. Accenture also points out that in many enterprise environments, fewer than 10% of users and workflows can account for the majority of AI consumption.
Financial services is already showing us what that growth curve might look like. RBC has reportedly built more than 200 AI models, while its LLM token consumption has increased more than 500% in a year.
So we are arriving at a somewhat inconvenient equation for enterprise technology leaders: token prices can fall dramatically while the total AI bill continues to climb.
If the price of something falls by 50%, but consumption rises by 500%, your CFO is probably not sending you a thank-you note.
And this is before agentic AI really begins to operate at scale.
Agents change the equation
A traditional chatbot interaction is relatively straightforward: A customer asks a question, the model processes some context and produces a response. An agent, on the other hand, may do something very different.
It might interpret the request, retrieve information from multiple sources, reason over that information, make a plan, call an internal system, evaluate the result, invoke another tool, retry a failed step, consult another model and finally validate that the intended task was actually completed.
From the customer's perspective, one thing happened. From the infrastructure's perspective, however, a lot happened. An agent might perform dozens of invisible actions to create one visible outcome.
The FinOps Foundation has highlighted that a single agent run that fills its context window with retrieved documents can cost orders of magnitude more than a simple query.
That makes one of the assumptions from the chatbot era increasingly dangerous: one customer interaction does not necessarily equal one AI interaction.
That is fantastic when we are talking about automation and productivity. It is worth understanding when we are talking about economics.
The token may not even be the right unit of measurement
This is where the conversation gets more interesting.
A case study cited by Innobu involving an AI-supported mortgage workflow at a regional bank found that tokens represented only 22% of the total AI cost per transaction. The rest came from things like tool calls, vector database queries, human review and compliance activities.
That finding should matter to anyone implementing AI inside a regulated institution.
A bank could negotiate an extraordinary price per million tokens and still build a very expensive AI process. Conversely, a more expensive model could actually produce better economics if it completes a task correctly on the first attempt, while a cheaper model requires repeated retries, larger context windows, additional orchestration and human intervention.
That is why I think the long-term conversation will move away from simply measuring cost per token and toward measuring cost per successful outcome. This measures, for example:Β
- What did it cost to resolve the servicing request?Β
- What did it cost to complete the mortgage task?Β
- What did it cost to investigate the transaction, execute the payment, identify the next best action or complete an onboarding step?Β
- What business value did that outcome create?
That is a much more useful definition of tokenomics. Otherwise, we risk becoming incredibly sophisticated at measuring the price of something without understanding whether it accomplished anything valuable.
Not every banking problem needs the biggest model
Enterprises are beginning to get much more thoughtful about model selection, marking a shift happening alongside the economics discussion.
During the first stage of generative AI adoption, there was an understandable tendency to equate model size with model quality. If the largest, smartest model performed best, the safest approach was often to use it.
But banking contains thousands of bounded, repeatable problems.
Classifying an intent, identifying a transaction, answering a payment question, retrieving a specific policy, mapping a servicing request or executing a known workflow does not necessarily require the most sophisticated reasoning model available.
You would not hire a neurosurgeon to take someone's temperature, and AI architecture is beginning to learn the same lesson.
JPMorgan CFO Jeremy Barnum has spoken publicly about the importance of using the right model for the right purpose and making sure the organization ultimately receives value from it. That mindset is becoming increasingly important because model selection is no longer simply an engineering decision. It is an economic decision.
That is also why I believe that the future probably is not one model doing everything. Instead, small language models, specialized models and intelligent model routing will become increasingly important inside enterprise AI architectures.
A sophisticated reasoning model may orchestrate a complex process while smaller, domain-specific models perform predictable tasks beneath it. A conversational interaction may use one model, a voice task another, an operational workflow another, and a highly complex reasoning problem a frontier model only when it is actually needed.
AI needs an economic architecture
This is where I think financial institutions have an opportunity to get ahead of the problem.
Banks spent much of the last decade developing greater discipline around cloud economics. FinOps emerged because organizations realized that infinitely scalable infrastructure could also create infinitely scalable invoices.
AI introduces another layer of variable consumption, but with an important difference: increasingly, the software itself can make decisions that generate additional consumption.
Agents can invoke models. Models can invoke tools. Agents can invoke other agents. Context windows can grow. Reasoning loops can multiply. Retrieval requests can expand. All of this can occur without a human explicitly clicking a button for each action.
That makes visibility and control much more important.
Consumption is not spread evenly. Accenture's finding that fewer than 10% of users and workflows can account for the majority of AI usage means a small number of workflows are probably driving most of the bill.
Financial institutions are going to need to understand:
- Which workflows are consuming AI resources
- Which models are serving those workflows
- How much context is being passed
- How many model and tool calls are happening behind an interaction
- Where retries are occurring
- Where a smaller model could perform the same task Β
- What business result all of that activity is producing
The goal should not be to consume the fewest tokens possible, but to use the right amount of intelligence, from the right model, for the right outcome.
If $5 of AI consumption eliminates a $25 servicing expense, improves customer experience and resolves a task instantly, that may be an excellent trade. If we spend $25 worth of inference, orchestration and human review to automate a $5 problem, we probably need to revisit the architecture.
From AI experimentation to AI economics
Nobody running a contact center wakes up obsessing over the cost of a second of telephone bandwidth. They care about cost per interaction, resolution rate, employee productivity and customer outcomes. AI will likely mature in much the same way.
Enterprise AI adoption started with asking whether we could build certain experiences. Then the question became whether we could govern them safely and responsibly. Increasingly, the next question will be whether we can operate them economically at scale.
Instead of these being competing priorities, I look at this as a sign of maturity.
The institutions that get this right will likely develop a much clearer understanding of cost per completed outcome, cost per resolved customer interaction, token consumption by workflow, model utilization by business process, human intervention per AI-assisted task and the value generated relative to inference cost.
Over time, I suspect we may even stop talking much about token prices themselves.
The banks that get tokenomics right will not necessarily be the ones consuming the fewest tokens.
They will be the ones that can answer a much more important question: We consumed all of those tokens. What did we get for them?




.png)