The new partner opportunities in AI tokenomics
MSPs and SIs could help customers manage costs, optimise workloads and make smarter architecture choices.
Tokenomics is becoming a core conversation organisations should be having, particularly as AI usage grows and businesses turn their attention to ROI and costs, according to Lizzy Jones, head of data and AI, Interactive.
As AI becomes embedded in workplace processes, managing token costs has quickly emerged as a new headache for organisations. For partners, the economics of tokens, or tokenomics, presents both opportunities and challenges.
The lack of consistency in how tokens are used and measured across different models makes costs difficult to predict.
“One use-case on one model may have a completely different profile on another model,” she said.
Jones believes tokenomics could see technology budgets managed by business units. Rather than IT carrying all AI costs centrally, each individual unit may have to account for their own usage.
MSPs, SIs and other technology partners could find new opportunities in AI cost management such as helping organisations establish usage thresholds, track consumption and design tightly controlled workloads.
“What’s really interesting to me is what becomes a managed service versus does a business want to retain that internally,” Jones said.
Why agentic AI can spike costs
Gartner predicts increased competition will reduce token costs by 95 percent by 2030; however, overall inference costs will rise as usage grows and there are more large-parameter models.
Agentic systems can consume tokens at a much higher rate than individual users. As these systems become more complex and handle more tasks, AI costs could spike. This was illustrated at the Dell Technologies Forum in Sydney in August.
John Roese, Dell Technologies global CTO, described an internal cost experiment, where two teams were tasked with solving the same customer configuration and pricing problem with agentic AI.
One team loaded all the files and documents into a million-token context window and let the model complete the work. It consumed 25 million tokens. The other team used an agent to coordinate existing configuration and pricing tools. It used just 20,000 tokens.
“Twenty five million and 20,000 tokens. That's the degree of difference in two different approaches architecturally to exactly the same problem and almost exactly the same outcome,” Roese said.
The experiment demonstrates why tokenomics isn’t simply a token-cost issue, but an architectural one that has the potential to create cost advantages or structural cost problems.
“Those workloads are not just the model and not just the agent. They’re the entire system around it,” he said.
For partners, there’s an opportunity to move beyond implementation to help customers design, reconfigure and optimise systems for AI cost and token usage.
Choosing the right model matters
Token bill shock is a real risk, according to David Keane, CEO and co-founder, SCX.ai, who suggested a large enterprise workloads might consume tens of billions of tokens a month, potentially resulting in a cost of up to $1 million a year.
One way to reduce costs is with open-weight models, where trained model weights are available for organisations to run on their own infrastructure or private cloud.
“If you use these open-weights models, you can reduce your cost by as much as 40 percent, and even go more than that if you do it really well,” Keane said.
Another approach is to use tools that direct workloads to the most appropriate model, something SCX.ai has developed with its ‘Router’ feature to assess tasks and models for cost and efficiency.
Keane sees an opportunity for partners to help manage AI and token consumption, but warns that they risk being bypassed if the large AI providers develop more of their own professional services.
“They’re going to come in and say, ‘We’ll do the implementation for you as well’,” he ended.