Tokenomics: Why Making AI Pay Is Tricky



A selection of AI apps on a phone

Free chat interfaces from giants like Microsoft and Google feel like pure luck, but behind the scenes each request is broken into tokens that are processed by deep neural networks. Those tokens are what the providers charge for in paid iterations of their services.



The big question is how to lock a price into a $1‑ask or a yearly ceiling when the number of tokens used is unpredictable. “Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn’t make any sense, honestly, because we don’t know,” comments Simon Gooch of Saviynt, who is experimenting with AI agents in its identity‑management stack.



Tokens are the fundamental building blocks of LLMs and agentic AI. A single prompt and its output are split into many tokens, each of which drives the computational work. The same input can generate a different number of tokens each time, and multiple agents can amplify usage so the total cost is hard to predict.



Goldman Sachs projects that token consumption will rise 24‑fold between 2026 and 2030, reaching 120 quadrillion tokens a month. Even as per‑token prices drop, the sheer volume of tokens needed by businesses makes the bill explode – a problem Microsoft and Uber have each faced when internal developers over‑used third‑party tools.



"People are finding it really hard to manage that cost… it’s a non‑deterministic output, so it’s a non‑deterministic value," says Will Venters of the London School of Economics.


Smaller firms adopt flat‑fee arrangements to keep token spend predictable, but big vendors warn that such pools may trigger stricter governance once shareholder pressure mounts. Meanwhile, companies are urged to tighten prompts, choose cheaper models, and plan for incremental costs when launching AI features to large user bases.



Bill Peterson of Sumo Logic notes that even with advanced AI offerings, businesses need to “pay by results” or bundle incidents to meet customer expectations, yet any chosen pricing scheme will be vulnerable to swings caused by the evolving cost structure of LLM providers.



Understanding tokenomics and securing a forecasted pricing model will become essential for firms to monetize AI responsibly, as the technology continues to permeate every industry.