Token Wars: Why Pricing AI Services is a Moving Target
Large language models (LLMs) run on a quirky unit of measurement called a token – a fragment of text that the model can parse and generate. When you type a prompt, the system splits it into thousands of tokens, runs them through complex maths, then stitches back a response also tokenised. Because of this choreography, the cost is directly tied to the number of tokens a request consumes.
The paradox is simple: token prices have plummeted over the past few years, yet the amount of tokenised data flowing through companies is exploding. Goldman Sachs estimates that token consumption could hit 120 quadrillion tokens a month by 2030, as firms shift from legacy software to “AI agent” stacks. For most organisations this means a lack of visibility: they may not realise how many tokens they are using until the bill lands.
Simon Gooch of Saviynt calls the issue a nightmare: “Trying to tie someone into a cost model for the next 12, 24, or 30 months makes little sense because we don’t know how many tokens will be consumed.” Companies are already seeing the consequences. Microsoft has reportedly throttled its engineers’ usage of third‑party coding assistants, and Uber reportedly exhausted its AI coding budget within months of launch.
Why is tokenisation so unpredictable? A seemingly trivial tweak to a prompt can change the answer, and the same question can elicit differing token counts depending on the model and the user’s context. When multiple agents are chained together – what many call agentic AI – the token budget multiplies further. Each decision, call or test adds a new layer of tokens that was not‑there in the previous iteration.
Experts are looking for control. Will Venters, an LSE professor, notes that “people are having a difficult time managing the cost because it’s a non‑deterministic output, so it’s a non‑deterministic value.” Some organisations are turning to flat‑fee personal accounts that slip under the radar, while others are experimenting with prompt‑engineering to slash unnecessary token use. The CFO of UK accounting software firm iplicit, Rob Steele, urges teams to treat prompts like a grocery list: “You wouldn’t send a friend to buy your weekly shop without knowing what you expect.”
Yet cost is an ever‑shifting baseline. Bill Peterson from Sumo Logic says the company is still negotiating how to price new, AI‑driven security services. “We’re having serious internal conversations,” he admits. The market is moving between bulk‑billing, results‑based pricing and incident bundles, but if AI providers change their own rates, any model can crumble overnight.
What does this mean for the future? As LLMs become central to products offered to thousands of users, token costs could balloon beyond what a fixed‑budget team can reconcile. Though the cost is hard to solder, many argue that better value arises from spending tokens wisely: more computation can yield richer outcomes. Still, companies will need to translate token usage into charge‑to‑customer fees that stay predictable for budgets.


















