If you’ve ever used ChatGPT for free, it feels like a steal.


Big firms like Microsoft, Google and Anthropic poured billions into building the AI they power, so a free request from ChatGPT or Claude costs almost nothing to the user.


But those companies still need to recoup the investment. They offer paid plans that unlock extra features such as code generation, advanced analytics or enterprise‑level security.


The real puzzle is how to price services that spin up a large language model for each request. Tokens, the tiny units of text the model tokenises, are unpredictable: a different prompt can use a vastly different number of tokens, and the same prompt can vary at different times.


Token usage is exploding. Goldman Sachs projects monthly token consumption could jump twenty‑four times by 2030, reaching about 120 quadrillion tokens a month as organisations move from single‑model use to multi‑agent workflows.


Simon Gooch from Saviynt explains that companies can’t lock in a cost model for years because “we don’t know” how many tokens they’ll use. Will Venters of LSE adds that staff experimenting with AI can burn tokens faster than the bill appears.


Companies are trying work‑arounds: using flat‑fee personal accounts to keep an eye on token use, tightening prompts for precision, or bundling token costs into overall product pricing. Bill Peterson from Sumo Logic says they’re still debating whether to raise flat prices, charge per result, or bundle incidents.


In short, token costs can be a calculator that keeps changing. But managers have to decide if the value they get from AI justifies the spend, and how best to pass that cost onto end‑users in a way that’s understandable and budget‑friendly.