Who Pays For The Tokens

Software used to be a seat. You paid per person per month, it went in the overhead line, and nobody thought about it again. AI tools are quietly moving to metered pricing, and tokens do not behave like seats.
OpenAI's business access to ChatGPT inside Microsoft Word is billed on token-based pricing at API rates rather than a flat fee. That is one product, but it is the direction, and it lands awkwardly in an industry that sells time.
Why A Retainer Cannot Absorb It
A retainer is a fixed price against a variable workload, and agencies have always managed that variance with people, who cost the same whether they are busy or not. Metered compute is the opposite: it costs nothing when idle and scales directly with how much work is being produced.
So the busiest month, which used to be the one where the retainer finally paid off, becomes the one where an unbudgeted line appears.
Scope creep used to cost hours. Now it also costs compute, and only one of those is in the contract.
Who Ends Up Paying For Tokens
Initially the agency, because the sums are small enough to swallow and raising them looks petty. That holds until a client asks for forty variants of everything, at which point the cost is visible and the conversation happens from a weak position.
The Clause Worth Adding Now
Not a markup on tokens, which clients will resist and which invites scrutiny of margin. A volume assumption: this scope assumes a stated number of generated variants, and materially more is a new conversation. It is the same principle as rounds of amends, applied to the thing that now actually scales.
More from The AI Shift.