The Rise and Fall of ‘Tokenmaxxing’: How Businesses Fell Out of Love With Burning Tokens

Tech’s biggest players once flaunted their AI token counts like a badge of honour. Now, businesses are learning that using AI more isn’t the same as using it well.

Meta, Microsoft, and Salesforce were amongst those boasting about the number of tokens – the unit of data used to power AI models – being consumed by their employees.

The name ‘tokenmaxxing’ – borrowed from internet slang to indicate maximising or ‘maxxing’ a part of your life – became a badge of honour for many engineers and a perceived proxy of productivity amongst some of the world’s largest tech companies. 

The Information reported that a Meta employee, now famously, created a leaderboard tracking how many tokens the company’s workforce use.

The leaderboard reportedly showed the company’s top 250 token users and awarded those using the most tokens with titles such as ‘Token Legend’ and ‘Cache Wizard’ given to leaders.

The leaderboard didn’t last. Once the story broke, Meta pulled it down, and with it went the whole ‘tokenmaxxing’ ethos. The public conversation flipped almost overnight. All of a sudden, heavy token usage stopped looking like a productivity flex and started looking like a cost problem.

From ‘Tokenmaxxing’ to ‘Tokenomics’

For businesses using large language models (LLMs) like ChatGPT and Claude, token pricing is typically based on how many tokens go into the model (input) – such as a prompt given to an AI model – and how many tokens come out (output) – such as the response given by the AI model. Different models charge different rates, and output tokens are often more expensive than input tokens.

ChatGPT launching in November 2022 had numerous consequences. As businesses scrambled to implement AI into their workflows, teams often went ‘all in’, with most companies prioritising speed over intentional implementation to remain competitive. This often meant that the cost of using AI was overlooked.

Companies can typically gauge how many tokens they’re using from having a grasp on the number of prompts being generated by their workforce everyday. 

For a frontier model like Claude Fable 5, a single short prompt costs roughly between $0.005 and $0.02, heavy code generation costs between $0.05 and $0.10, and a deep document analysis costs between $0.25 and $0.50. 

This may seem inexpensive, but in a large enterprise with 100,000 people, if everyone did a single short prompt, it would cost between $500 and $2000. When it comes to running AI agents, a study from data platform Splunk found that agents use up five to 30 times more tokens than prompting alone.

When businesses recognised using token consumption as a productivity tracker is costly, and isn’t a good indicator of productivity as more prompts don’t necessarily mean better outcomes, sentiment quickly shifted from ‘tokenmaxxing’ to token efficiency.

From Unstructured Prompting to Intentional AI

A big element of shifting to token efficiency for many businesses, has been getting away from ‘unstructured prompting’ towards bespoke, intentional AI placement and usage.

The shift to efficiency has, for many companies, meant abandoning “unstructured prompting” in favour of deliberate, purpose-built AI deployment.

Tim Ringel, Global CEO of agency Meet The People, argues that rising costs come down to LLM providers getting paid once through token fees, and again through the ad revenue generated by giving consumers access to their models. The result, in his view, is that “token spend has skyrocketed and companies are pulling back on non-essential AI” use among staff. He claims that Meet The People’s response has been a deliberately cautious rollout – figuring out who actually needs which tool, and building in-house systems that guide employees toward structured, purposeful use rather than “unstructured prompting.” Ringel says the strategy has paid off on both fronts: the agency’s headcount has kept growing year over year, right alongside its revenue and profits.

A similar logic is playing out in marketing departments, according to Kenny Johnson, Principal Product Manager at Cloudflare, where AI is already delivering faster campaign analysis, more ad copy variants, and in-house image and video production. But the businesses getting real value, he argues, aren’t the ones reaching for the most powerful model by default – “a campaign performance summary doesn’t need the same model as a customer-facing content generator.” His advice is to match the model to the job and back it with an actual budget, so AI usage scales the way any other business cost would. As he puts it: “Match the model to the task, put a real budget behind it, and you get AI that scales like any other business tool.”

There’s a related fix on offer from Mateusz Rumiński, VP of Product at ad tech company PrimeAudience: shifting workloads toward Retrieval-Augmented Generation (RAG). Agentic pipelines, RAG, and multi-step processes all draw on far more compute than a single prompt, he notes, and companies make that worse by defaulting to the newest, flashiest model even when a cheaper one would do the job just as well. The real obstacle, he says, is that “it’s difficult to judge where the difference in quality becomes meaningful.”

These accounts highlight that ‘tokenmaxxing’ era was never really about productivity, it was about visibility – a number that was easy to track and easier to show off. Efficiency is a less flattering metric, but a far more honest one. As the novelty of AI adoption wears off and the invoices keep arriving, the companies pulling ahead look less like the ones using AI the most, and more like the ones using it deliberately.

Subscribe to our newsletter for updates

Join thousands of media and marketing professionals by signing up for our newsletter.

"*" indicates required fields

This field is for validation purposes and should be left unchanged.

Share

Related Posts

Popular Articles

Featured Posts

Menu